multimodal
972 个项目 · ⭐ 708.1k[ICML 2026] GRACE-VLM: deployable INT4 Qwen3-VL via quantization-aware distillation.
[SIGIR 2022] Hybrid Transformer with Multi-level Fusion for Multimodal Knowledge Graph Completion
Self-hostable multimodal chat with local LLMs (Ollama/OpenAI): PDF RAG, image chat, and Whisper voice, Streamlit + Docker.
MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
A Practical Course on Embeddings, RAG, Multimodal Models, and Agents with Amazon Nova.
Video2Music: Suitable Music Generation from Videos using an Affective Multimodal Transformer model
A toolkit for building computer use AI agents
[ACM TOMM 2023] - Composed Image Retrieval using Contrastive Learning and Task-oriented CLIP-based Features
The official code of "VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning" [NeurIPS25]
共 972 条 · 第 14 / 49 页