[ICML'26] Toward Human-like Audio-Visual Intelligence of Omni-MLLMs
[ICLR 2026🔥] SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
Open Translator: Speech To Speech and Speech to text Translator with voice cloning and other cool features
This repo contains the original implementation of VAuLT, the Vision-and-Augmented-Language Transformer. We provide instructions to download some multimodal social-media datasets, and scripts to experiment with. VAuLT is a stack of Transformers, a LM like BERT that preprocesses the text input of ViLT
A Comprehensive Speech Processing Algorithms Library for research and production use
DiscordNPC lets you interact with ChatGPT through a Discord voice channel, enabling a natural conversation.
Listen4Me bot converts audio messages into text.
An OpenAI Compatible API which integrates LLM, Embedding and Reranker. 一个集成 LLM、Embedding 和 Reranker 的 OpenAI 兼容 API
共 9101 条 · 第 432 / 456 页