speech
141 个项目 · ⭐ 372.2k🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
SoftVC VITS Singing Voice Conversion
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools
kaldi-asr/kaldi is the official location of the Kaldi project.
Build voice agents with open-source models
AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
:robot: :speech_balloon: Deep learning for Text to Speech (Discussion forum: https://discourse.mozilla.org/c/tts)
Silero VAD: pre-trained enterprise-grade Voice Activity Detector
ModelScope: bring the notion of Model-as-a-Service to life.
EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine
Officially maintained, supported by PaddlePaddle, including CV, NLP, Speech, Rec, TS, big models and so on.
Silero Models: pre-trained text-to-speech models made embarrassingly simple
Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
Voice Recognition to Text Tool / 一个离线运行的本地音视频转字幕工具,输出json、srt字幕、纯文字格式
Noise supression using deep filtering
共 141 条 · 第 1 / 8 页