vqa
23 个项目 · ⭐ 19.0kA modular framework for vision & language multimodal research from Facebook AI Research (FAIR)
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, SAM, interactive image editing, etc. Try it at igpt.opengvlab.com (支持DragGAN、ChatGPT、ImageBind、SAM的在线Demo系统)
streamline the fine-tuning process for multimodal models: PaliGemma 2, Florence-2, and Qwen2.5-VL
[ICLR'24] Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
OmniFusion — a multimodal model to communicate using text and images
mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video (ICML 2023)
[CVPR 2025] Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
[CVPR 2026 Highlight] ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
🚁 Can Vision-Language Models Think from the Sky? UAVReason for Aerial Reasoning and Generation
共 23 条 · 第 1 / 2 页