mllm
55 个项目 · ⭐ 72.1kLarge-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
Agent S: an open agentic framework that uses computers like a human
Mobile-Agent: The Powerful GUI Agent Family
From Chain-of-Thought prompting to OpenAI o1 and DeepSeek-R1 🍓
Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
Eagle: Frontier Vision-Language Models with Data-Centric Strategies
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
A family of lightweight multimodal models.
OpenEMMA, a permissively licensed open source "reproduction" of Waymo’s EMMA model.
NEO Series: Native Vision-Language Models from First Principles
Personal Project: MPP-Qwen14B & MPP-Qwen-Next(Multimodal Pipeline Parallel based on Qwen-LM). Support [video/image/multi-image] {sft/conversations}. Don't let the poverty limit your imagination! Train your own 8B/14B LLaVA-training-like MLLM on RTX3090/4090 24GB.
[ECCV2024] Grounded Multimodal Large Language Model with Localized Visual Tokenization
🏭 Mega Scale Multimodal DataPipeline for SOTA Foundation Models
Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Pre-training Dataset and Benchmarks
共 55 条 · 第 1 / 3 页