vision-language-model
97 个项目 · ⭐ 117.0kTurn documents into AI-ready Markdown with visual understanding
Evaluating text-to-image/video/3D models with VQAScore
[ECCV2024] Grounded Multimodal Large Language Model with Localized Visual Tokenization
Flame is an open-source multimodal AI system designed to translate UI design mockups into high-quality React code. It leverages vision-language modeling, automated data synthesis, and structured training workflows to bridge the gap between design and front-end development.
[ ICLR 2024 ] Official Codebase for "InstructCV: Instruction-Tuned Text-to-Image Diffusion Models as Vision Generalists"
An open-source implementation for training LLaVA-NeXT.
Stream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across various modality combinations.
A Simple, Lightweight, and Extensible Serving Framework for X-AnyLabeling
Seamlessly integrate state-of-the-art transformer models into robotics stacks
共 97 条 · 第 2 / 5 页