Generate editable scientific SVG figures from method text with local SAM3 and dual-provider routing.
WhisperCrabs is a simple terminal-based floating recording button, click, and transcribe to input
jupyter notebooks to fine tune whisper models on Vietnamese using Colab and/or Kaggle and/or AWS EC2
Mandarin Chinese audio datasets aligned with Montreal Forced Aligner
Many ASRs under one roof. With Benchmarking... answering the question. What is the best ASR for my dataset?
Using Snowboy and Google Cloud speech api in Electron for voice recognition
This repository contains a short introduction on the topic of audio and speech processing -- from basics to applications.
Speech-to-text from videos and audios (including youtube and tiktok links)
共 26086 条 · 第 1256 / 1305 页