llava
54 个项目 · ⭐ 67.4kVisual Instruction Tuning for Qwen2 Base Model
A bug-free and improved implementation of LLaVA-UHD, based on the code from the official repo
Learn how multimodal AI merges text, image, and audio for smarter models
A project to show howto use SpringAI with Ollama to chat with the documents in a library. Documents are stored in a normal/vector database. The AI is used to create embeddings from documents that are stored in the vector database. The vector database is used to query for the nearest document. That document is used by the AI to generate the answer.
[ICLR 2026] MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding
Llava, Ollama and Streamlit | Create POWERFUL Image Analyzer Chatbot for FREE - Windows & Mac
[ICLR 2026🔥] SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
[WACV 2025] I Dream My Painting: Connecting MLLMs and Diffusion Models via Prompt Generation for Text-Guided Multi-Mask Inpainting
Multimodal document QA: vision + retrieval over PDFs (LLaVA + LlamaIndex)
共 54 条 · 第 3 / 3 页