multimodal
972 个项目 · ⭐ 707.9kA novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
A curated list of Multimodal Related Research.
Offline inference engine for art, real-time voice conversations, LLM powered chatbots and automated workflows
Xtreme1 is an all-in-one data labeling and annotation platform for multimodal data training and supports 3D LiDAR point cloud, image, and LLM.
Implementation of CoCa, Contrastive Captioners are Image-Text Foundation Models, in Pytorch
Turn documents into AI-ready Markdown with visual understanding
Unifying 3D Mesh Generation with Language Models
The Self-Coding System for Your App — Alan AI SDK for Cordova
MOVA: Towards Scalable and Synchronized Video–Audio Generation
Codebase for Aria - an Open Multimodal Native MoE
🩺 首个会看胸部X光片的中文多模态医学大模型 | The first Chinese Medical Multimodal Model that Chest Radiographs Summarization.
共 972 条 · 第 5 / 49 页