multimodal
972 个项目 · ⭐ 707.8kOpen-source framework for building agentic apps in JavaScript, Go, Dart, and Python, built and used in production by Google
A framework for efficient model inference with omni-modality models
Solve Visual Understanding with Reinforced VLMs
A visual playground for agentic workflows: Iterate over your agents 10x faster
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines
A modular framework for vision & language multimodal research from Facebook AI Research (FAIR)
A Next-Generation Training Engine Built for Ultra-Large MoE Models
Align Anything: Training All-modality Model with Feedback
Curated tutorials and resources for Large Language Models, AI Painting, and more.
Easily turn large sets of image urls to an image dataset. Can download, resize and package 100M urls in 20h on one machine.
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
Fengshenbang-LM(封神榜大模型)是IDEA研究院认知计算与自然语言研究中心主导的大模型开源体系,成为中文AIGC和认知智能的基础设施。
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenarios, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS.
OpenMMLab Pre-training Toolbox and Benchmark
🪩 Create Disco Diffusion artworks in one line
SimpleMem: Efficient Lifelong Memory for LLM Agents — Text & Multimodal
共 972 条 · 第 2 / 49 页