speech-synthesis

42 projects share this GitHub topic

speech-synthesis — Speech ★18.4kspeech-synthesisDeepLearningExamples — ★14.8kspeech-to-speech — ★13kedge-tts — ★11.5kvits — ★7.9kespeak-ng — ★6.7kStyleTTS2 — ★6.3kabogen — ★5.8kRealtimeTTS — ★4kawesome-speech-recognition-speech-synthesis-papers — ★3.1ktacotron — ★3kMARS5-TTS — ★2.8kmarytts — ★2.6kchat-with-gpt — ★2.4kTacotron-2 — ★2.3kVieNeu-TTS — ★2.2kWaveRNN — ★2.2kSpeechT5 — ★1.4kartyom.js — ★1.3kIrene-Voice-Assistant — ★1.1kIrodori-TTS — ★1klocal-talking-llm — ★878glow-tts — ★712vits2 — ★642tiktok-voice — ★607Speech-Backbones — ★604StyleTTS — ★465Ming-UniAudio — ★450speech-recognition-uk — ★439ProDiff — ★432FastDiff — ★423nnmnkwii — ★399Freeze-Omni — ★396Stream-Omni — ★392MsEdgeTTS — ★335GenerSpeech — ★333VocGAN — ★321voiceai — ★306manim-voiceover — ★305axiom-voice-agent — ★144Qwen3-TTS-EasyFinetuning — ★121DeepLearningExamples★ 14.8kspeech-to-speech★ 13kedge-tts★ 11.5kvits★ 7.9kespeak-ng★ 6.7kStyleTTS2★ 6.3kabogen★ 5.8kRealtimeTTS★ 4kawesome-speech-recogniti…★ 3.1ktacotron★ 3kMARS5-TTS★ 2.8kmarytts★ 2.6kchat-with-gpt★ 2.4kTacotron-2★ 2.3kVieNeu-TTS★ 2.2kWaveRNN★ 2.2kSpeechT5★ 1.4kartyom.js★ 1.3kIrene-Voice-Assistant★ 1.1kIrodori-TTS★ 1klocal-talking-llm★ 878glow-tts★ 712vits2★ 642tiktok-voice★ 607Speech-Backbones★ 604StyleTTS★ 465Ming-UniAudio★ 450speech-recognition-uk★ 439ProDiff★ 432FastDiff★ 423nnmnkwii★ 399Freeze-Omni★ 396Stream-Omni★ 392MsEdgeTTS★ 335GenerSpeech★ 333VocGAN★ 321voiceai★ 306manim-voiceover★ 305axiom-voice-agent★ 144Qwen3-TTS-EasyFinetuning★ 121

Lines connect members that are measurably related to each other. Dot size reflects stars.

🧬 Members
Speech
A scalable generative AI framework built for researchers and developers working on Large Language Models,…
★ 18.4k
DeepLearningExamples
State-of-the-Art Deep Learning scripts organized by models - easy to train and deploy with reproducible…
★ 14.8k
speech-to-speech
Build voice agents with open-source models
★ 13k
edge-tts
Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or…
★ 11.5k
vits
VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
★ 7.9k
espeak-ng
eSpeak NG is an open source speech synthesizer that supports more than hundred languages and accents.
★ 6.7k
StyleTTS2
StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large…
★ 6.3k
abogen
Generate audiobooks from EPUBs, PDFs and text with synchronized captions.
★ 5.8k
RealtimeTTS
Converts text to speech in realtime
★ 4k
awesome-speech-recognition-speech-synthesis-papers
Automatic Speech Recognition (ASR), Speaker Verification, Speech Synthesis, Text-to-Speech (TTS), Language…
★ 3.1k
tacotron
A TensorFlow implementation of Google's Tacotron speech synthesis with pre-trained model (unofficial)
★ 3k
MARS5-TTS
MARS5 speech model (TTS) from CAMB.AI
★ 2.8k
marytts
MARY TTS -- an open-source, multilingual text-to-speech synthesis system written in pure java
★ 2.6k
chat-with-gpt
An open-source ChatGPT app with a voice
★ 2.4k
Tacotron-2
DeepMind's Tacotron-2 Tensorflow implementation
★ 2.3k
VieNeu-TTS
Vietnamese TTS with instant voice cloning • On-device • Real-time CPU inference • 24kHz audio quality…
★ 2.2k
WaveRNN
WaveRNN Vocoder + TTS
★ 2.2k
SpeechT5
Unified-Modal Speech-Text Pre-Training for Spoken Language Processing
★ 1.4k
artyom.js
A voice control - voice commands - speech recognition and speech synthesis javascript library. Create your…
★ 1.3k
Irene-Voice-Assistant
Ирина - русский голосовой ассистент для работы оффлайн.…
★ 1.1k
Irodori-TTS
A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control
★ 1k
local-talking-llm
A talking LLM that runs on your own computer without needing the internet.
★ 878
glow-tts
A Generative Flow for Text-to-Speech via Monotonic Alignment Search
★ 712
vits2
VITS2: Improving Quality and Efficiency of Single-Stage Text-to-Speech with Adversarial Learning and…
★ 642
tiktok-voice
Simple Python script to interact with the TikTok TTS API
★ 607
Speech-Backbones
This is the main repository of open-sourced speech technology by Huawei Noah's Ark Lab.
★ 604
StyleTTS
Official Implementation of StyleTTS
★ 465
Ming-UniAudio
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
★ 450
speech-recognition-uk
🇺🇦 Speech Recognition & Synthesis for Ukrainian
★ 439
ProDiff
PyTorch Implementation of ProDiff (ACM-MM'22) with a Extremely-Fast diffusion speech synthesis pipeline
★ 432
FastDiff
PyTorch Implementation of FastDiff (IJCAI'22)
★ 423
nnmnkwii
Library to build speech synthesis systems designed for easy and fast prototyping.
★ 399
Freeze-Omni
✨✨Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
★ 396
Stream-Omni
Stream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across…
★ 392
MsEdgeTTS
A simple Azure Speech Service module that uses the Microsoft Edge Read Aloud API.…
★ 335
GenerSpeech
PyTorch Implementation of GenerSpeech (NeurIPS'22): a text-to-speech model towards zero-shot style transfer…
★ 333
VocGAN
VocGAN: A High-Fidelity Real-time Vocoder with a Hierarchically-nested Adversarial Network
★ 321
voiceai
Set of 📝 with 🔗 to help those building Voice AI agents 🎙️🤖
★ 306
manim-voiceover
Manim plugin for all things voiceover
★ 305
axiom-voice-agent
Run a <400ms latency Voice Agent on just 4GB VRAM. Fully offline, no API keys required. Optimized for GTX…
★ 144
Qwen3-TTS-EasyFinetuning
Easy fine-tuning for Qwen3-TTS: Fast voice cloning and high-quality multilingual speech synthesis.
★ 121
Core-AI-Framework-Lab
A practical lab for exploring Apple's Core AI framework, model assets, specialization, and on-device…
★ 64
🔗 Related families

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.