speech-to-text
986 个项目 · ⭐ 742.1kOpen source, local voice-to-text (Wispr Flow alternative)
A dockerfile to run deepspeech-server
Coqui STT offline engine API for NodeJs developers. With a simple HTTP ASR server.
Automatic Speech Recognition Dataset for Oromo Language
A collection of useful tools for handling speech recognition data
Docker image for a self-hosted WhisperLive real-time speech-to-text server, powered by faster-whisper. Provides WebSocket streaming for live audio transcription and an OpenAI-compatible REST API. Supports all Whisper models, VAD, NVIDIA GPU (CUDA) acceleration, offline mode, and multi-arch (amd64, arm64).
ComfyUI nodes for Qwen3-ASR (0.6B/1.7B) and ForcedAligner. Supports high-accuracy ASR and language identification for 52 languages/dialects, including 22 Chinese dialects and various English accents. Features word-level timestamps, long audio transcription, and VRAM-optimized inference.
A lightweight library for normalizing speech transcripts before computing WER
共 986 条 · 第 37 / 50 页