multimodal
972 个项目 · ⭐ 707.9kAn Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
The Self-Coding System for Your App — Alan AI SDK for React Native
An MCP Multimodal AI Agent with eyes and ears!
Port of MiniGPT4 in C++ (4bit, 5bit, 6bit, 8bit, 16bit CPU inference with GGML)
CLIP inference in plain C/C++ with no extra dependencies
Flame is an open-source multimodal AI system designed to translate UI design mockups into high-quality React code. It leverages vision-language modeling, automated data synthesis, and structured training workflows to bridge the gap between design and front-end development.
A hub for various industry-specific schemas to be used with VLMs.
Official implementation for "Break-A-Scene: Extracting Multiple Concepts from a Single Image" [SIGGRAPH Asia 2023]
Official code of "EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model"
A simple "Be My Eyes" web app with a llama.cpp/llava backend
VisualWebArena is a benchmark for multimodal agents.
[NeurIPS 2025 Spotlight] Official implementation of Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
共 972 条 · 第 8 / 49 页