text-to-speech

76 projects share this GitHub topic

text-to-speech — MoneyPrinterTurbo ★119.1ktext-to-speechunsloth — ★75.4kGPT-SoVITS — ★59.9kOpenMontage — ★55kChatTTS — ★39.8kOpenVoice — ★37kindex-tts — ★22kdia — ★19.3kpyvideotrans — ★18.4kedge-tts — ★11.5kvits — ★7.9kStyleTTS2 — ★6.3kabogen — ★5.8kdograh — ★5.5kODS — ★5.4kfastrtc — ★4.6kMOSS-TTS — ★4.1kRealtimeTTS — ★4kTTS-WebUI — ★3.2kelevenlabs-python — ★3kvall-e — ★3kMARS5-TTS — ★2.8kmarytts — ★2.6kChatTTS_colab — ★2.6kvall-e — ★2.2kopenai-edge-tts — ★2.1kparlor — ★2kAuto-Synced-Translated-Dubs — ★1.7kGenie-TTS — ★1.7kvllm-mlx — ★1.6kRCLI — ★1.5kOuteTTS — ★1.4kSpeech-AI-Forge — ★1.4kmlx-tune — ★1.4kGPA — ★1.3ksoprano — ★1.2kFoundation-Models-Framework-Lab — ★1.2kXZVoice — ★1.2kdia2 — ★1.2kPatter — ★1kbotium-speech-processing — ★943unsloth★ 75.4kGPT-SoVITS★ 59.9kOpenMontage★ 55kChatTTS★ 39.8kOpenVoice★ 37kindex-tts★ 22kdia★ 19.3kpyvideotrans★ 18.4kedge-tts★ 11.5kvits★ 7.9kStyleTTS2★ 6.3kabogen★ 5.8kdograh★ 5.5kODS★ 5.4kfastrtc★ 4.6kMOSS-TTS★ 4.1kRealtimeTTS★ 4kTTS-WebUI★ 3.2kelevenlabs-python★ 3kvall-e★ 3kMARS5-TTS★ 2.8kmarytts★ 2.6kChatTTS_colab★ 2.6kvall-e★ 2.2kopenai-edge-tts★ 2.1kparlor★ 2kAuto-Synced-Translated-D…★ 1.7kGenie-TTS★ 1.7kvllm-mlx★ 1.6kRCLI★ 1.5kOuteTTS★ 1.4kSpeech-AI-Forge★ 1.4kmlx-tune★ 1.4kGPA★ 1.3ksoprano★ 1.2kFoundation-Models-Framew…★ 1.2kXZVoice★ 1.2kdia2★ 1.2kPatter★ 1kbotium-speech-processing★ 943

Lines connect members that are measurably related to each other. Dot size reflects stars.

🧬 Members
MoneyPrinterTurbo
利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD…
★ 119.1k
unsloth
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma…
★ 75.4k
GPT-SoVITS
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
★ 59.9k
OpenMontage
World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent…
★ 55k
ChatTTS
A generative speech model for daily dialogue.
★ 39.8k
OpenVoice
Instant voice cloning by MIT and MyShell. Audio foundation model.
★ 37k
index-tts
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
★ 22k
dia
A TTS model capable of generating ultra-realistic dialogue in one pass.
★ 19.3k
pyvideotrans
Translate the video from one language to another and embed dubbing & subtitles.
★ 18.4k
edge-tts
Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or…
★ 11.5k
vits
VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
★ 7.9k
StyleTTS2
StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large…
★ 6.3k
abogen
Generate audiobooks from EPUBs, PDFs and text with synchronized captions.
★ 5.8k
dograh
Open source voice AI platform. Self-hosted alternative to Vapi and Retell. On Prem, BYOK across Speech to…
★ 5.5k
ODS
Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG,…
★ 5.4k
fastrtc
The python library for real-time communication
★ 4.6k
MOSS-TTS
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS…
★ 4.1k
RealtimeTTS
Converts text to speech in realtime
★ 4k
TTS-WebUI
A single Gradio + React WebUI with extensions for ACE-Step, OmniVoice, Kimi Audio, Piper TTS, GPT-SoVITS,…
★ 3.2k
elevenlabs-python
The official Python SDK for the ElevenLabs API.
★ 3k
vall-e
An unofficial PyTorch implementation of the audio LM VALL-E
★ 3k
MARS5-TTS
MARS5 speech model (TTS) from CAMB.AI
★ 2.8k
marytts
MARY TTS -- an open-source, multilingual text-to-speech synthesis system written in pure java
★ 2.6k
ChatTTS_colab
🚀 一键部署(含离线整合包)!基于 ChatTTS ,支持流式输出、音色抽卡、长音频生…
★ 2.6k
vall-e
PyTorch implementation of VALL-E(Zero-Shot Text-To-Speech), Reproduced Demo…
★ 2.2k
openai-edge-tts
Free, high-quality text-to-speech API endpoint to replace OpenAI, Azure, or ElevenLabs
★ 2.1k
parlor
On-device, real-time multimodal AI with features similar to GPT-Live
★ 2k
Auto-Synced-Translated-Dubs
Automatically translates the text of a video based on a subtitle file, and then uses AI voice services to…
★ 1.7k
Genie-TTS
GPT-SoVITS ONNX Inference Engine & Model Converter
★ 1.7k
vllm-mlx
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX,…
★ 1.6k
RCLI
Talk to your Mac, query your docs, no cloud required. On-device voice AI + RAG
★ 1.5k
OuteTTS
Interface for OuteTTS models.
★ 1.4k
Speech-AI-Forge
🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a…
★ 1.4k
mlx-tune
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR…
★ 1.4k
GPA
[AutoArk] GPA (General Purpose Audio) can do ASR, TTS and voice conversion with one tiny model!
★ 1.3k
soprano
Soprano: Instant, Ultra-Realistic Text-to-Speech
★ 1.2k
Foundation-Models-Framework-Lab
A practical lab for building, testing, and evaluating apps with Apple's Foundation Models framework.
★ 1.2k
XZVoice
Free and open source text-to-speech software
★ 1.2k
dia2
TTS model capable of streaming conversational audio in realtime.
★ 1.2k
Patter
Open-source voice-AI SDK. The Vapi/Retell alternative for builders who want to own the stack. Give your AI…
★ 1k
botium-speech-processing
Botium Speech Processing
★ 943
bark.cpp
Suno AI's Bark model in C/C++ for fast text-to-speech generation
★ 866
june
Local voice chatbot for engaging conversations, powered by Ollama, Hugging Face Transformers, and Coqui TTS…
★ 787
CloneTTS
A lightweight, offline Android Text-to-Speech (TTS) engine enabling seamless system-wide voice cloning and…
★ 738
glow-tts
A Generative Flow for Text-to-Speech via Monotonic Alignment Search
★ 712
bark-voice-cloning-HuBERT-quantizer
The code for the bark-voicecloning model. Training and inference.
★ 711
voicebox-pytorch
Implementation of Voicebox, new SOTA Text-to-speech network from MetaAI, in Pytorch
★ 698
ai-video-editor
Open-source, local-first video editor where creators and AI agents edit the same real timeline.
★ 666
LLaSA_training
LLaSA: Scaling Train-time and Inference-time Compute for LLaMA-based Speech Synthesis
★ 660
f5-tts-mlx
Implementation of F5-TTS in MLX
★ 640
examples
★ 617
tiktok-voice
Simple Python script to interact with the TikTok TTS API
★ 607
Awesome-LLMs-meet-Multimodal-Generation
🔥🔥🔥 A curated list of papers on LLMs-based multimodal generation (image, video, 3D and audio).
★ 552
vits2_pytorch
unofficial vits2-TTS implementation in pytorch
★ 548
orpheus-tts-local
Run Orpheus 3B Locally With LM Studio
★ 544
e2-tts-pytorch
Implementation of E2-TTS, "Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS", in Pytorch
★ 516
vixtts-demo
A Vietnamese Voice Cloning Text-to-Speech Model ✨
★ 515
ComfyUI-OmniVoice-TTS
OmniVoice TTS nodes for ComfyUI - Zero-shot multilingual text-to-speech with voice cloning, voice design, and…
★ 509
google-speech-v2
:speech_balloon: Reverse Engineering Google's Speech To Text API (v2)
★ 469
StyleTTS
Official Implementation of StyleTTS
★ 465
ai-avatar-system
🎭 AI Avatar / digital human platform — upload a photo, clone a voice, talk to any face in real time with…
★ 451
dectalk
Modern builds for the 90s/00s DECtalk text-to-speech application.
★ 451
ProDiff
PyTorch Implementation of ProDiff (ACM-MM'22) with a Extremely-Fast diffusion speech synthesis pipeline
★ 432
FastDiff
PyTorch Implementation of FastDiff (IJCAI'22)
★ 423
Cross-Lingual-Voice-Cloning
Tacotron 2 - PyTorch implementation with faster-than-realtime inference modified to enable cross lingual…
★ 359
easevoice-trainer
EaseVoice Trainer is a simple and user-friendly voice cloning and speech model trainer.
★ 352
StreamingKokoroJS
Unlimited text-to-speech in the Browser using Kokoro-JS, 100% local, 100% open source
★ 346
Speech-and-Text
Speech to text (PocketSphinx, Iflytex API, Baidu API) and text to speech (pyttsx3) |…
★ 342
sdk
AI video generation SDK — JSX for videos. One API for Kling, Flux, ElevenLabs, Veed. Built on Vercel AI SDK.
★ 335
Whisper-TikTok
From AI tools to TikTok video creation using FFMPEG, Microsoft Edge read aloud and OpenAI Whisper model
★ 335
gemini-youtube-automation
A fully autonomous AI Agent/Python pipeline that utilizes Large Language Models (LLMs) like Gemini to…
★ 334
voiceai
Set of 📝 with 🔗 to help those building Voice AI agents 🎙️🤖
★ 306
self-hosted-ai-stack
Deploy a complete self-hosted AI stack with Docker Compose: Ollama, LiteLLM, AnythingLLM, Whisper,…
★ 141
Qwen3-TTS-EasyFinetuning
Easy fine-tuning for Qwen3-TTS: Fast voice cloning and high-quality multilingual speech synthesis.
★ 121
BlueTTS
Fastest Open Source TTS Model
★ 86
Core-AI-Framework-Lab
A practical lab for exploring Apple's Core AI framework, model assets, specialization, and on-device…
★ 64
🔗 Related families

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.