← 返回专题广场
large-multimodal-models
27 个项目 · ⭐ 8.5k1
22 天前
最近推送
2
2024-10-10
最近推送
3
A Framework of Small-scale Large Multimodal Models
Python
⭐ 1.0k
⑂ 103
Apache-2.0
· 2026-07-23推送
2026-07-23
最近推送
4
25 天前
最近推送
5
2025-06-29
最近推送
6
An open-source implementation for training LLaVA-NeXT.
Python
⭐ 440
⑂ 23
· 2024-10-23推送
2024-10-23
最近推送
7
2024-08-24
最近推送
8
Stream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across various modality combinations.
Python
⭐ 391
⑂ 45
GPL-3.0
· 2025-06-17推送
2025-06-17
最近推送
9
2026-02-28
最近推送
11
[CVPR 2026] LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Python
⭐ 260
⑂ 16
Apache-2.0
· 2026-06-24推送
2026-06-24
最近推送
12
Official implementation of GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
Python
⭐ 255
⑂ 19
Apache-2.0
· 2025-05-05推送
2025-05-05
最近推送
13
2025-09-20
最近推送
14
2024-09-26
最近推送
15
Embed arbitrary modalities (images, audio, documents, etc) into large language models.
Python
⭐ 189
⑂ 16
Apache-2.0
· 2024-03-27推送
2024-03-27
最近推送
16
[CVPR 2026] OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe
Python
⭐ 166
⑂ 5
Apache-2.0
· 2026-03-30推送
2026-03-30
最近推送
17
2026-05-09
最近推送
18
2025-07-09
最近推送
19
2024-10-19
最近推送
20
PhyX: Does Your Model Have the "Wits" for Physical Reasoning?
Python
⭐ 55
⑂ 1
MIT
· 2026-03-16推送
2026-03-16
最近推送
共 27 条 · 第 1 / 2 页