WhisperCrabs is a simple terminal-based floating recording button, click, and transcribe to input
jupyter notebooks to fine tune whisper models on Vietnamese using Colab and/or Kaggle and/or AWS EC2
Mandarin Chinese audio datasets aligned with Montreal Forced Aligner
Many ASRs under one roof. With Benchmarking... answering the question. What is the best ASR for my dataset?
Using Snowboy and Google Cloud speech api in Electron for voice recognition
This repository contains a short introduction on the topic of audio and speech processing -- from basics to applications.
Speech-to-text from videos and audios (including youtube and tiktok links)
Free, open-source, offline, safe and secure AI Cantonese transcription, in your device.
A project combining roguelike with LLMs, RAG, Text2Speech, and Speech2Text
共 40708 条 · 第 1972 / 2036 页