← 返回专题广场
vision-language
29 个项目 · ⭐ 31.2k21
2025-12-01
最近推送
22
[CVPR 2024] The official implementation of paper "synthesize, diagnose, and optimize: towards fine-grained vision-language understanding"
Jupyter Notebook
⭐ 52
⑂ 0
· 2025-06-16推送
2025-06-16
最近推送
24
VL-JEPA inspired pipeline — compress images/text locally via Ollama, send compact payloads to any LLM API. Cut token costs by ~80%.
Python
⭐ 24
⑂ 1
NOASSERTION
· 2026-07-16推送
2026-07-16
最近推送
25
[ICLR 2026] - Spectral Concept Selection and Cross-modal Representation Learning for Generalized Category Discovery
Python
⭐ 23
⑂ 0
MIT
· 2026-03-18推送
2026-03-18
最近推送
26
2023-10-07
最近推送
27
Multimodal document QA: vision + retrieval over PDFs (LLaVA + LlamaIndex)
Python
⭐ 13
⑂ 0
NOASSERTION
· 2026-06-03推送
2026-06-03
最近推送
28
ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval
Jupyter Notebook
⭐ 11
⑂ 2
MIT
· 2026-03-10推送
2026-03-10
最近推送
29
Vision-Language Models on AMD GPUs — LLaVA, MiniGPT-4, Idefics on ROCm 🚀
Python
⭐ 11
⑂ 0
MIT
· 2026-06-24推送
2026-06-24
最近推送
共 29 条 · 第 2 / 2 页