multimodal
972 个项目 · ⭐ 707.9kFrom Chain-of-Thought prompting to OpenAI o1 and DeepSeek-R1 🍓
Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
MTEB: State-of-the-art evaluation of embeddings across languages and modalities
InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, SAM, interactive image editing, etc. Try it at igpt.opengvlab.com (支持DragGAN、ChatGPT、ImageBind、SAM的在线Demo系统)
Represent, send, store and search multimodal data
[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure
Easily compute clip embeddings and build a clip retrieval system with them
Images to inference with no labeling (use foundation models to train supervised models).
streamline the fine-tuning process for multimodal models: PaliGemma 2, Florence-2, and Qwen2.5-VL
[EMNLP-2024] Build multimodal language agents for fast prototype and production
Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
mPLUG-Owl: The Powerful Multi-modal Large Language Model Family
HuixiangDou: Overcoming Group Chat Scenarios with LLM-based Technical Assistance
共 972 条 · 第 3 / 49 页