multimodal
972 个项目 · ⭐ 708.1kDemocratization of "PaLI: A Jointly-Scaled Multilingual Language-Image Model"
Neural Machine Translation with universal Visual Representation (ICLR 2020)
An open source chat bot architecture for voice/vision (and multimodal) assistants, local(CPU/GPU bound) and remote(I/O bound) to run.
Multi-modal speech separation task data generation script on LRS3 data set.
Robust multimodal integration method implemented in PyTorch and TensorFlow
Multimodal AI agent, an interactive data studio with on-demand ML inference, media generation, and a database explore
Unofficial implementation and experiments related to Set-of-Mark (SoM) 👁️
Deep Research through Multi-Agents, using GraphRAG
Open-source multimodal dataset: 98K+ clinical cases & 139K+ medical images from PubMed Central
共 972 条 · 第 21 / 49 页