asr
215 个项目 · ⭐ 227.7kComfyUI nodes for Qwen3-ASR (0.6B/1.7B) and ForcedAligner. Supports high-accuracy ASR and language identification for 52 languages/dialects, including 22 Chinese dialects and various English accents. Features word-level timestamps, long audio transcription, and VRAM-optimized inference.
A lightweight library for normalizing speech transcripts before computing WER
基于 SenseVoice 的 Windows 本地语音转文字工具,支持 OpenAI 格式 API 润色,低延迟,高精度。
PAFTS : Library That Preprocessing Audio For TTS.
Scripts for training Kaldi for German speech recognition (ASR).
Easy to use Multi-Provider ASR/Speech To Text and NLP engine
Streaming ASR, audio captioning and speech QA over a whisper-style encoder with pluggable LLM backends
This is an OpenAI Whisper automatic speech recognition microservice
Keras(Tensorflow) implementations of Automatic Speech Recognition
Presenting Collection of Pretrained Models. Links to pretrained models in NLP and voice.
SDKs and docs for Skit's speech to text service
Local-first, capability-aware, traceable audio and video transcription Skill for Codex
An implementation for "Conformer: Convolution-augmented Transformer for Speech Recognition" Paper
共 215 条 · 第 10 / 11 页