speech-recognition
551 个项目 · ⭐ 697.6kA collection of useful tools for handling speech recognition data
Docker image for a self-hosted WhisperLive real-time speech-to-text server, powered by faster-whisper. Provides WebSocket streaming for live audio transcription and an OpenAI-compatible REST API. Supports all Whisper models, VAD, NVIDIA GPU (CUDA) acceleration, offline mode, and multi-arch (amd64, arm64).
ComfyUI nodes for Qwen3-ASR (0.6B/1.7B) and ForcedAligner. Supports high-accuracy ASR and language identification for 52 languages/dialects, including 22 Chinese dialects and various English accents. Features word-level timestamps, long audio transcription, and VRAM-optimized inference.
Speech to Text with self-supervised learning based on wav2vec 2.0 framework using Hugging Face's Transformer
Convert images or audio files to plain text on the command line
Voice memos recorded from the microphone, transcribed offline to text and converted to Joplin notes
A collection of utilities for handling IPA phones.
Ear is a desktop app that will help you transcribe what is playing on your computer!
Scripts for training Kaldi for German speech recognition (ASR).
Offline streaming speech-to-text in the browser
共 551 条 · 第 22 / 28 页