← 返回专题广场
vision-language-pretraining
6 个项目 · ⭐ 29.6k1
Janus-Series: Unified Multimodal Understanding and Generation Models
Python
⭐ 17.8k
⑂ 2.2k
MIT
· 2025-02-01推送
2025-02-01
最近推送
2
LAVIS - A One-stop Library for Language-Vision Intelligence
Jupyter Notebook
⭐ 11.3k
⑂ 1.1k
BSD-3-Clause
· 2026-06-03推送
2026-06-03
最近推送
3
Official Repository of paper VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
Python
⭐ 293
⑂ 19
CC-BY-4.0
· 2025-08-05推送
2025-08-05
最近推送
4
[ICLR 2026] Uni-CoT: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
Python
⭐ 236
⑂ 7
Apache-2.0
· 2026-05-31推送
2026-05-31
最近推送
5
Easy wrapper for inserting LoRA layers in CLIP.
Python
⭐ 40
⑂ 2
MIT
· 2024-06-17推送
2024-06-17
最近推送
6
Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator
Python
⭐ 34
⑂ 1
Apache-2.0
· 2026-04-15推送
2026-04-15
最近推送