vision-language-model
97 个项目 · ⭐ 117.0kVision Document Retrieval (ViDoRe): Benchmark. Evaluation code for the ColPali paper.
Archived snapshot of Thinking-with-Visual-Primitives
From scratch implementation of a vision language model in pure PyTorch
[NeurIPS-2023] Annual Conference on Neural Information Processing Systems
[ICML 2026] GRACE-VLM: deployable INT4 Qwen3-VL via quantization-aware distillation.
Embed arbitrary modalities (images, audio, documents, etc) into large language models.
[NeurIPS 2024] CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
[TMLR 25] SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
NEWTON: Agentic Planning for Physically Grounded Video Generation
[ICLR2025 Oral] ChartMoE: Mixture of Diversely Aligned Expert Connector for Chart Understanding
共 97 条 · 第 3 / 5 页