mllm
55 个项目 · ⭐ 72.1k[CVPR 2026] LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
A CPU Realtime VLM in 500M. Surpassed Moondream2 and SmolVLM. Training from scratch with ease.
Official code for Paper "Mantis: Multi-Image Instruction Tuning" [TMLR 2024 Best Paper]
mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video (ICML 2023)
[ICCV25 Oral] Token Activation Map to Visually Explain Multimodal LLMs
[CVPR 2026] OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe
中文医学多模态大模型 Large Chinese Language-and-Vision Assistant for BioMedicine
[NeurIPS'25 Spotlight] Official implementation of "JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation"
[NeurIPS 2025] Deep Memory Backtracking for Long Video Understanding
R1-Track: Direct Application of MLLMs to Visual Object Tracking via Reinforcement Learning.
[CVPR 2025] Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
在千问最新的多模态image-text模型Qwen3-VL-4B-Instruct 进行多种lora微调对比效果,通过langchain+RAG+多智能体(Multi-Agent)进行部署
ElasticMM: Elastic and Efficient MLLM Serving System
共 55 条 · 第 2 / 3 页