← 返回专题广场
visual-language-learning
8 个项目 · ⭐ 36.1k1
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
Python
⭐ 25.0k
⑂ 2.8k
Apache-2.0
· 2024-08-12推送
2024-08-12
最近推送
2
Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
Python
⭐ 3.6k
⑂ 359
BSD-3-Clause
· 2025-05-13推送
2025-05-13
最近推送
3
2024-03-05
最近推送
4
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Python
⭐ 2.9k
⑂ 174
Apache-2.0
· 2025-05-26推送
2025-05-26
最近推送
5
An open-source implementation for training LLaVA-NeXT.
Python
⭐ 440
⑂ 23
· 2024-10-23推送
2024-10-23
最近推送
6
2024-09-11
最近推送
7
(AAAI 2024) BLIVA: A Simple Multimodal LLM for Better Handling of Text-rich Visual Questions
Python
⭐ 261
⑂ 25
BSD-3-Clause
· 2024-04-15推送
2024-04-15
最近推送