multi-modality
22 个项目 · ⭐ 53.2k[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
🏄 Scalable embedding, reasoning, ranking for images and sentences with CLIP
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
An open-source implementation for training LLaVA-NeXT.
Effortless plugin and play Optimizer to cut model training costs by 50%. New optimizer that is 2x faster than Adam on LLMs.
MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
Embed arbitrary modalities (images, audio, documents, etc) into large language models.
An open-source cloud-native of large multi-modal models (LMMs) serving framework.
An all-new Language Model That Processes Ultra-Long Sequences of 100,000+ Ultra-Fast
Seed, Code, Harvest: Grow Your Own App with Tree of Thoughts!
My implementation of Kosmos2.5 from the paper: "KOSMOS-2.5: A Multimodal Literate Model"
CrossCLR: Cross-modal Contrastive Learning For Multi-modal Video Representations, ICCV 2021
共 22 条 · 第 1 / 2 页