speech-to-text

59 projects share this GitHub topic

speech-to-text — whisper.cpp ★51.8kspeech-to-textmeetily — ★30.2kllamafile — ★25.8kwhisperX — ★23.1kscreenpipe — ★21.3kFunASR — ★20.1kpyvideotrans — ★18.4kspeech-to-speech — ★13kspeech_recognition — ★9kannyang — ★6.8kFunClip — ★6.2kdograh — ★5.5kODS — ★5.4kwhisper-jax — ★4.7kstt — ★4.7kfastrtc — ★4.6kopenless — ★3.4kBayLing-Speech — ★3.1kpluely — ★2.6kWhisperJAV — ★2.2ktensorflow-speech-recognition — ★2.2kvllm-mlx — ★1.6kRCLI — ★1.5kminutes — ★1.5kmlx-tune — ★1.4kwhisper-ctranslate2 — ★1.3kgp.nvim — ★1.3kartyom.js — ★1.3kmuesli — ★1.1kPatter — ★1kVoiceStreamAI — ★958botium-speech-processing — ★943react-speech-recognition — ★842whisper-playground — ★833whisper_mic — ★788june — ★787whisper.unity — ★749awesome-large-audio-models — ★737curses — ★711speech-demo — ★710expo-speech-recognition — ★652meetily★ 30.2kllamafile★ 25.8kwhisperX★ 23.1kscreenpipe★ 21.3kFunASR★ 20.1kpyvideotrans★ 18.4kspeech-to-speech★ 13kspeech_recognition★ 9kannyang★ 6.8kFunClip★ 6.2kdograh★ 5.5kODS★ 5.4kwhisper-jax★ 4.7kstt★ 4.7kfastrtc★ 4.6kopenless★ 3.4kBayLing-Speech★ 3.1kpluely★ 2.6kWhisperJAV★ 2.2ktensorflow-speech-recogn…★ 2.2kvllm-mlx★ 1.6kRCLI★ 1.5kminutes★ 1.5kmlx-tune★ 1.4kwhisper-ctranslate2★ 1.3kgp.nvim★ 1.3kartyom.js★ 1.3kmuesli★ 1.1kPatter★ 1kVoiceStreamAI★ 958botium-speech-processing★ 943react-speech-recognition★ 842whisper-playground★ 833whisper_mic★ 788june★ 787whisper.unity★ 749awesome-large-audio-mode…★ 737curses★ 711speech-demo★ 710expo-speech-recognition★ 652

Lines connect members that are measurably related to each other. Dot size reflects stars.

🧬 Members
whisper.cpp
Port of OpenAI's Whisper model in C/C++
★ 51.8k
meetily
Privacy first, AI meeting assistant with 4x faster Parakeet/Whisper live transcription, speaker diarization,…
★ 30.2k
llamafile
Distribute and run LLMs with a single file.
★ 25.8k
whisperX
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
★ 23.1k
screenpipe
YC (S26) | Open Computer History | Record your screen continuously locally and provide context to your agents…
★ 21.3k
FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker…
★ 20.1k
pyvideotrans
Translate the video from one language to another and embed dubbing & subtitles.
★ 18.4k
speech-to-speech
Build voice agents with open-source models
★ 13k
speech_recognition
Speech recognition module for Python, supporting several engines and APIs, online and offline.
★ 9k
annyang
💬 Speech recognition for your site
★ 6.8k
FunClip
FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio…
★ 6.2k
dograh
Open source voice AI platform. Self-hosted alternative to Vapi and Retell. On Prem, BYOK across Speech to…
★ 5.5k
ODS
Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG,…
★ 5.4k
whisper-jax
JAX implementation of OpenAI's Whisper model for up to 70x speed-up on TPU.
★ 4.7k
stt
★ 4.7k
fastrtc
The python library for real-time communication
★ 4.6k
openless
Hold a key, speak, release — AI-polished text appears at your cursor in any app. Open-source voice input…
★ 3.4k
BayLing-Speech
LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon…
★ 3.1k
pluely
The Open Source Alternative to Cluely - A lightning-fast, privacy-first AI assistant that works seamlessly…
★ 2.6k
WhisperJAV
ASR/STT subtitle generator. Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD. Noise-robust for JAV
★ 2.2k
tensorflow-speech-recognition
🎙Speech recognition using the tensorflow deep learning framework, sequence-to-sequence neural networks
★ 2.2k
vllm-mlx
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX,…
★ 1.6k
RCLI
Talk to your Mac, query your docs, no cloud required. On-device voice AI + RAG
★ 1.5k
minutes
Every meeting, every idea, every voice note, searchable by your AI. Open-source, privacy-first conversation…
★ 1.5k
mlx-tune
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR…
★ 1.4k
whisper-ctranslate2
Whisper command line client compatible with original OpenAI client based on CTranslate2.
★ 1.3k
gp.nvim
Gp.nvim (GPT prompt) Neovim AI plugin: ChatGPT sessions & Instructable text/code operations & Speech to text…
★ 1.3k
artyom.js
A voice control - voice commands - speech recognition and speech synthesis javascript library. Create your…
★ 1.3k
muesli
Muesli: agent-native local meeting transcription + dictation for macOS (Granola + WisprFlow alternative)
★ 1.1k
Patter
Open-source voice-AI SDK. The Vapi/Retell alternative for builders who want to own the stack. Give your AI…
★ 1k
VoiceStreamAI
Near-Realtime audio transcription using self-hosted Whisper and WebSocket in Python/JS
★ 958
botium-speech-processing
Botium Speech Processing
★ 943
react-speech-recognition
💬Speech recognition for your React app
★ 842
whisper-playground
Build real time speech2text web apps using OpenAI's Whisper https://openai.com/blog/whisper/
★ 833
whisper_mic
Project that allows one to use a microphone with OpenAI whisper.
★ 788
june
Local voice chatbot for engaging conversations, powered by Ollama, Hugging Face Transformers, and Coqui TTS…
★ 787
whisper.unity
Running speech to text model (whisper.cpp) in Unity3d on your local machine.
★ 749
awesome-large-audio-models
Collection of resources on the applications of Large Language Models (LLMs) in Audio AI.
★ 737
curses
Speech to Text and KB input captions for OBS, VRChat, Twitch chat and Discord
★ 711
speech-demo
语音api示例
★ 710
expo-speech-recognition
Speech Recognition for React Native Expo projects
★ 652
speech-to-text
Real-time transcription using faster-whisper
★ 613
WhisperS2T
An Optimized Speech-to-Text Pipeline for the Whisper Model Supporting Multiple Inference Engine
★ 577
Awesome-Korean-Speech-Recognition
한국어 음성인식 STT API 리스트. 각 성능 벤치마크.
★ 534
vosk-browser
A speech recognition library running in the browser thanks to a WebAssembly build of Vosk
★ 526
ai-avatar-system
🎭 AI Avatar / digital human platform — upload a photo, clone a voice, talk to any face in real time with…
★ 451
speech-recognition-uk
🇺🇦 Speech Recognition & Synthesis for Ukrainian
★ 439
whisper-youtube
🔉 Youtube Videos Transcription with OpenAI's Whisper
★ 421
VRCT
VRCT(VRChat Chatbox Translator & Transcription)
★ 406
PreenCut
AI-Powered Video Retrieval & Clipping Tool
★ 405
Stream-Omni
Stream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across…
★ 392
self-supervised-speech-recognition
speech to text with self-supervised learning based on wav2vec 2.0 framework
★ 380
qvac
QVAC - Local AI SDK and libraries for building private, cross-platform, peer-to-peer AI applications. Run…
★ 356
insanely-fast-whisper-api
An API to transcribe audio with OpenAI's Whisper Large v3!
★ 354
Speech-and-Text
Speech to text (PocketSphinx, Iflytex API, Baidu API) and text to speech (pyttsx3) |…
★ 342
MsEdgeTTS
A simple Azure Speech Service module that uses the Microsoft Edge Read Aloud API.…
★ 335
watch-skill
Video understanding and self-verification for AI agents. Turn videos, streams, and agent screen recordings…
★ 324
voiceai
Set of 📝 with 🔗 to help those building Voice AI agents 🎙️🤖
★ 306
self-hosted-ai-stack
Deploy a complete self-hosted AI stack with Docker Compose: Ollama, LiteLLM, AnythingLLM, Whisper,…
★ 141
🔗 Related families

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.