multimodal-large-language-models
55 个项目 · ⭐ 35.0kMobile-Agent: The Powerful GUI Agent Family
StarVector is a foundation model for SVG generation that transforms vectorization into a code generation task. Using a vision-language modeling architecture, StarVector processes both visual and textual inputs to produce high-quality SVG code with remarkable precision.
A cross-platform video structuring (video analysis) framework based on CV models & mLLM.
mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
A family of lightweight multimodal models.
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
NEO Series: Native Vision-Language Models from First Principles
Personal Project: MPP-Qwen14B & MPP-Qwen-Next(Multimodal Pipeline Parallel based on Qwen-LM). Support [video/image/multi-image] {sft/conversations}. Don't let the poverty limit your imagination! Train your own 8B/14B LLaVA-training-like MLLM on RTX3090/4090 24GB.
(Accepted by IJCV) Liquid: Language Models are Scalable and Unified Multi-modal Generators
Official code of "EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model"
[NeurIPS 2025 Spotlight] Official implementation of Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
[ICLR 2026] VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLM
共 55 条 · 第 1 / 3 页