multimodal
972 个项目 · ⭐ 708.0kGemini Vision & Image Generation MCP for Claude Desktop and Claude Code
[ICLR 2025] Official code repository for "TULIP: Token-length Upgraded CLIP"
🦀 Rust powered LLM, Whisper, Embedding inference, backed by 🤗 candle from HuggingFace
Learn how multimodal AI merges text, image, and audio for smarter models
Framework for processing and filtering datasets
Generalized cross-modal NNs; new audiovisual benchmark (IEEE TNNLS 2019)
A unified Model Context Protocol server for MiniMax CLI (mmx)
Free unlimited Gemini API for any OpenAI-compatible app.
A prototype user experience concept for building interactive worlds and telling stories at the same time by sketching and speaking; Ph.D. Thesis Project and ACM UIST Publication: "DrawTalking: Building Interactive Worlds by Sketching and Speaking"; Paper: https://dl.acm.org/doi/10.1145/3654777.3676334
Un framework in Italiano ed Inglese, che permette di chattare con i propri documenti in RAG, anche multimediali (audio, video, immagini e OCR). It is an Italian and English framework, which allows you to chat with your documents in RAG, including multimedia (audio, video, images and OCR).
A beautiful Gemini-powered chat interface with multimodal support
共 972 条 · 第 32 / 49 页