audio
145 个项目 · ⭐ 772.1kmusic genre classification : LSTM vs Transformer
📣 Find sentiments, tags, entities, and actions in your voice recordings instantly
ASRecognition: just an easy-to-use library for Automatic Speech Recognition.
Work in progress - BBC News Labs digital paper edit project - React Client
Generate degraded speech datasets for noise-robust ASR benchmarking
Automatically generate multi-language subtitles using AWS AI/ML services. Machine generated subtitles can be edited to improve accuracy and downstream tracks will automatically be regenerated based on the edits. Built on Media Insights Engine (https://github.com/awslabs/aws-media-insights-engine)
real time japanese speech recognition translator using wav2vec2
Code and models for the paper "2D3MF: Deepfake Detection using Multi Modal Middle Fusion"
how to use the Google Cloud Speech API to transcribe audio/video files.
🎙️ Fast CLI tool to transcribe audio/video files to SRT format using OpenAI Whisper API
Docker image for a self-hosted WhisperLive real-time speech-to-text server, powered by faster-whisper. Provides WebSocket streaming for live audio transcription and an OpenAI-compatible REST API. Supports all Whisper models, VAD, NVIDIA GPU (CUDA) acceleration, offline mode, and multi-arch (amd64, arm64).
Streaming ASR, audio captioning and speech QA over a whisper-style encoder with pluggable LLM backends
共 145 条 · 第 7 / 8 页