audio
145 个项目 · ⭐ 772.1kPython API & command-line tool to easily transcribe speech-based video files into clean text
收集有关so-vits-svc、TTS、SD、LLMs的各种模型、应用以及文字、声音、图片、视频有关的model。
A collection of Audio and Speech pre-trained models.
OFASys: A Multi-Modal Multi-Task Learning System for Building Generalist Models
Timething is a library for aligning text transcripts with their audio recordings.
Your one-stop solution for voice dataset creation
PyTorch code for “TVLT: Textless Vision-Language Transformer” (NeurIPS 2022 Oral)
Implementation of a multimodal diffusion transformer in Pytorch
Transcribe your audio to text with this serverless component
Docker image for a self-hosted Whisper speech-to-text server with speaker diarization and OpenAI-compatible transcription and translation APIs. Powered by faster-whisper. Supports all Whisper models, NVIDIA GPU (CUDA) acceleration, JSON/SRT/VTT output, SSE streaming, offline mode, and multi-arch (amd64, arm64).
An open source chat bot architecture for voice/vision (and multimodal) assistants, local(CPU/GPU bound) and remote(I/O bound) to run.
Agent-first CLI for audio/video transcription via Whisper
A WebRTC-native, audio-first conversational-AI framework for Go.
Repository for the paper "Combining audio control and style transfer using latent diffusion", accepted at ISMIR 2024
共 145 条 · 第 6 / 8 页