multimodal-large-language-models
55 个项目 · ⭐ 35.0kYouku-mPLUG: A 10 Million Large-scale Chinese Video-Language Pre-training Dataset and Benchmarks
[CVPR 2026] LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
From scratch implementation of a vision language model in pure PyTorch
Official implementation of GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
[CVPR 2026] OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe
[ICLR 2025] AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
A multimodal chat interface with many tools.
[ICCV 2025] Explore the Limits of Omni-modal Pretraining at Scale
[ICLR2025 Oral] ChartMoE: Mixture of Diversely Aligned Expert Connector for Chart Understanding
[NeurIPS'25 Spotlight] Official implementation of "JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation"
共 55 条 · 第 2 / 3 页