multi-modal
49 个项目 · ⭐ 153.5kBuild and run agents you can see, understand and trust.
A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
Open-source framework for conversational voice AI agents
[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
ModelScope: bring the notion of Model-as-a-Service to life.
AI suite powered by state-of-the-art models and providing advanced AI/AGI functions. Includes AI personas, AGI functions, world-class Beam multi-model chats, text-to-image, voice, response streaming, code highlighting and execution, PDF import, presets for developers, much more. Deploy on-prem or in the cloud.
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.
Implementation / replication of DALL-E, OpenAI's Text to Image Transformer, in Pytorch
[EMNLP 2022] An Open Toolkit for Knowledge Graph Extraction and Construction
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
A C#/.NET library to run LLM (🦙LLaMA/LLaVA) on your local device efficiently.
Represent, send, store and search multimodal data
Project Page for "LISA: Reasoning Segmentation via Large Language Model"
[NeurIPS 2023] MotionGPT: Human Motion as a Foreign Language, a unified motion-language generation model using LLMs
Efficient Retrieval Augmentation and Generation Framework
共 49 条 · 第 1 / 3 页