multimodal
972 个项目 · ⭐ 707.9kAn official implementation for "CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval"
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
Resource, examples & tutorials for multimodal AI, RAG and agents using vector search and LLMs
Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.
SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.
RAG Time: A 5-week Learning Journey to Mastering RAG
NEO Series: Native Vision-Language Models from First Principles
Multimodal RL training framework for diffusion & omni models
library supporting NLP and CV research on scientific papers
Train Models Contrastively in Pytorch
共 972 条 · 第 6 / 49 页