speech-synthesis

99 Projekte teilen dieses GitHub-Topic

speech-synthesis — TTS ★45.8kspeech-synthesisVoxCPM — ★34.4kSpeech — ★17.8kleon — ★17.4kDeepLearningExamples — ★14.8ksupertonic — ★13.5kPaddleSpeech — ★12.7kedge-tts — ★11.6kvoice-pro — ★11.3kAmphion — ★10kespnet — ★9.9kso-vits-svc-fork — ★9.3kEmotiVoice — ★8.5kvits — ★7.9kmlx-audio — ★7.6kspeech-to-speech — ★7kespeak-ng — ★6.7kStyleTTS2 — ★6.3ksilero-models — ★6kabogen — ★5.4kDiffSinger — ★4.8kmetavoice-src — ★4.2kRealtimeTTS — ★4kTensorFlowTTS — ★4kawesome-speech-recognition-speech-synthesis-papers — ★3.1ktacotron — ★3klingvo — ★2.9kMARS5-TTS — ★2.8kmarytts — ★2.6khifi-gan — ★2.4kchat-with-gpt — ★2.4kTacotron-2 — ★2.3kVieNeu-TTS — ★2.2kIMS-Toucan — ★2.2kWaveRNN — ★2.2kdeepvoice3_pytorch — ★2kRHVoice — ★1.8kkalliope — ★1.8kParallelWaveGAN — ★1.6kdsnote — ★1.6kComfyUI_Custom_Nodes_AlekPet — ★1.5kVoxCPM★ 34.4kSpeech★ 17.8kleon★ 17.4kDeepLearningExamples★ 14.8ksupertonic★ 13.5kPaddleSpeech★ 12.7kedge-tts★ 11.6kvoice-pro★ 11.3kAmphion★ 10kespnet★ 9.9kso-vits-svc-fork★ 9.3kEmotiVoice★ 8.5kvits★ 7.9kmlx-audio★ 7.6kspeech-to-speech★ 7kespeak-ng★ 6.7kStyleTTS2★ 6.3ksilero-models★ 6kabogen★ 5.4kDiffSinger★ 4.8kmetavoice-src★ 4.2kRealtimeTTS★ 4kTensorFlowTTS★ 4kawesome-speech-recogniti…★ 3.1ktacotron★ 3klingvo★ 2.9kMARS5-TTS★ 2.8kmarytts★ 2.6khifi-gan★ 2.4kchat-with-gpt★ 2.4kTacotron-2★ 2.3kVieNeu-TTS★ 2.2kIMS-Toucan★ 2.2kWaveRNN★ 2.2kdeepvoice3_pytorch★ 2kRHVoice★ 1.8kkalliope★ 1.8kParallelWaveGAN★ 1.6kdsnote★ 1.6kComfyUI_Custom_Nodes_Ale…★ 1.5k

Linien verbinden Mitglieder, die messbar miteinander verwandt sind. Die Punktgröße spiegelt die Sterne wider.

🧬 Mitglieder
TTS
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
★ 45.8k
VoxCPM
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life…
★ 34.4k
Speech
A scalable generative AI framework built for researchers and developers working on Large Language Models,…
★ 17.8k
leon
🧠 Leon is your open-source personal assistant.
★ 17.4k
DeepLearningExamples
State-of-the-Art Deep Learning scripts organized by models - easy to train and deploy with reproducible…
★ 14.8k
supertonic
Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.
★ 13.5k
PaddleSpeech
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation,…
★ 12.7k
edge-tts
Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or…
★ 11.6k
voice-pro
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning…
★ 11.3k
Amphion
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support…
★ 10k
espnet
End-to-End Speech Processing Toolkit
★ 9.9k
so-vits-svc-fork
so-vits-svc fork with realtime support, improved interface and more features.
★ 9.3k
EmotiVoice
EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine
★ 8.5k
vits
VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
★ 7.9k
mlx-audio
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX…
★ 7.6k
speech-to-speech
Build local voice agents with open-source models
★ 7k
espeak-ng
eSpeak NG is an open source speech synthesizer that supports more than hundred languages and accents.
★ 6.7k
StyleTTS2
StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large…
★ 6.3k
silero-models
Silero Models: pre-trained text-to-speech models made embarrassingly simple
★ 6k
abogen
Generate audiobooks from EPUBs, PDFs and text with synchronized captions.
★ 5.4k
DiffSinger
DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism (SVS & TTS); AAAI 2022; Official code
★ 4.8k
metavoice-src
Foundational model for human-like, expressive TTS
★ 4.2k
RealtimeTTS
Converts text to speech in realtime
★ 4k
TensorFlowTTS
:stuck_out_tongue_closed_eyes: TensorFlowTTS: Real-Time State-of-the-art Speech Synthesis for Tensorflow 2…
★ 4k
awesome-speech-recognition-speech-synthesis-papers
Automatic Speech Recognition (ASR), Speaker Verification, Speech Synthesis, Text-to-Speech (TTS), Language…
★ 3.1k
tacotron
A TensorFlow implementation of Google's Tacotron speech synthesis with pre-trained model (unofficial)
★ 3k
lingvo
Lingvo
★ 2.9k
MARS5-TTS
MARS5 speech model (TTS) from CAMB.AI
★ 2.8k
marytts
MARY TTS -- an open-source, multilingual text-to-speech synthesis system written in pure java
★ 2.6k
hifi-gan
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
★ 2.4k
chat-with-gpt
An open-source ChatGPT app with a voice
★ 2.4k
Tacotron-2
DeepMind's Tacotron-2 Tensorflow implementation
★ 2.3k
VieNeu-TTS
Vietnamese TTS with instant voice cloning • On-device • Real-time CPU inference • 24kHz audio quality…
★ 2.2k
IMS-Toucan
Controllable and fast Text-to-Speech for over 7000 languages!
★ 2.2k
WaveRNN
WaveRNN Vocoder + TTS
★ 2.2k
deepvoice3_pytorch
PyTorch implementation of convolutional neural networks-based text-to-speech synthesis models
★ 2k
RHVoice
a free and open source speech synthesizer for Russian and other languages
★ 1.8k
kalliope
Kalliope is a framework that will help you to create your own personal assistant.
★ 1.8k
ParallelWaveGAN
Unofficial Parallel WaveGAN (+ MelGAN & Multi-band MelGAN & HiFi-GAN & StyleMelGAN) with Pytorch
★ 1.6k
dsnote
Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and…
★ 1.6k
ComfyUI_Custom_Nodes_AlekPet
Custom nodes that extend the capabilities of Comfyui
★ 1.5k
SpeechT5
Unified-Modal Speech-Text Pre-Training for Spoken Language Processing
★ 1.4k
open-speech-corpora
💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
★ 1.4k
Chatterbox-TTS-Server
Self-host the powerful Chatterbox TTS model. This server offers a user-friendly Web UI, flexible API…
★ 1.4k
merlin
This is now the official location of the Merlin project.
★ 1.3k
StreamSpeech
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech…
★ 1.3k
artyom.js
A voice control - voice commands - speech recognition and speech synthesis javascript library. Create your…
★ 1.3k
Irene-Voice-Assistant
Ирина - русский голосовой ассистент для работы оффлайн.…
★ 1.1k
AI-Waifu-Vtuber
AI Vtuber for Streaming on Youtube/Twitch
★ 1.1k
Irodori-TTS
A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control
★ 1.1k
Cognitive-Speech-TTS
Microsoft Text-to-Speech API sample code in several languages, part of Cognitive Services.
★ 1k
athena
an open-source implementation of sequence-to-sequence based speech processing engine
★ 968
NISQA
NISQA - Non-Intrusive Speech Quality and TTS Naturalness Assessment
★ 963
FireRedTTS
An Open-Sourced LLM-empowered Foundation TTS System
★ 909
diffwave
DiffWave is a fast, high-quality neural vocoder and waveform synthesizer.
★ 885
local-talking-llm
A talking LLM that runs on your own computer without needing the internet.
★ 883
NaturalVoiceSAPIAdapter
Make Azure natural TTS voices accessible to any SAPI 5-compatible application.
★ 871
alexandria-audiobook
AI-powered multi-voice audiobook generator — LLM script annotation, voice cloning, voice design, LoRA…
★ 840
vui
Real-time voice assistant — WebRTC streaming, faster-whisper ASR, local LLM, Vui Nano (300M) TTS. OpenAI…
★ 737
Confucius4-TTS
Confucius4-TTS: a Multilingual and Cross-Lingual Zero-Shot TTS Engine
★ 718
glow-tts
A Generative Flow for Text-to-Speech via Monotonic Alignment Search
★ 712
INTERSPEECH-2023-24-Papers
INTERSPEECH 2023-2024 Papers: A complete collection of influential and exciting research papers from the…
★ 685
vits2
VITS2: Improving Quality and Efficiency of Single-Stage Text-to-Speech with Adversarial Learning and…
★ 643
tiktok-voice
Simple Python script to interact with the TikTok TTS API
★ 609
Speech-Backbones
This is the main repository of open-sourced speech technology by Huawei Noah's Ark Lab.
★ 604
java-speech-api
The J.A.R.V.I.S. Speech API is designed to be simple and efficient, using the speech engines created by…
★ 542
CleanS2S
High-quality and streaming Speech-to-Speech interactive agent in a single file. …
★ 534
aspeak
A simple text-to-speech client for Azure TTS API.
★ 498
libfaceid
libfaceid is a research framework for prototyping of face recognition solutions. It seamlessly integrates…
★ 496
Awesome-Singing-Voice-Synthesis-and-Singing-Voice-Conversion
A paper and project list about the cutting edge Speech Synthesis, Text-to-Speech (TTS), Singing Voice…
★ 487
VibeVoiceFusion
VibeVoiceFusion is a full-stack, multi-speaker voice generation web system featuring LoRA fine-tuning, batch…
★ 484
StyleTTS
Official Implementation of StyleTTS
★ 466
speech_dataset
The dataset of Speech Recognition
★ 464
Ming-UniAudio
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
★ 451
echogarden
Cross-platform speech toolset, used from the command-line or as a Node.js library. Includes a variety of…
★ 444
speech-recognition-uk
🇺🇦 Speech Recognition & Synthesis for Ukrainian
★ 440
ProDiff
PyTorch Implementation of ProDiff (ACM-MM'22) with a Extremely-Fast diffusion speech synthesis pipeline
★ 432
FastDiff
PyTorch Implementation of FastDiff (IJCAI'22)
★ 424
WaveGrad
Implementation of WaveGrad high-fidelity vocoder from Google Brain in PyTorch.
★ 409
awesome-russian-speech
Russian speech technology links
★ 405
nnmnkwii
Library to build speech synthesis systems designed for easy and fast prototyping.
★ 399
Stream-Omni
Stream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across…
★ 390
Freeze-Omni
✨✨Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
★ 388
VoiceFlow-TTS
[ICASSP 2024] This is the official code for "VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching"
★ 376
UTMOSv2
UTokyo-SaruLab MOS Prediction System
★ 357
Dia-TTS-Server
Self-host the powerful Dia TTS model. This server offers a user-friendly Web UI, flexible API endpoints…
★ 352
DiffGAN-TTS
PyTorch Implementation of DiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion…
★ 349
PortaSpeech
PyTorch Implementation of PortaSpeech: Portable and High-Quality Generative Text-to-Speech
★ 342
MsEdgeTTS
A simple Azure Speech Service module that uses the Microsoft Edge Read Aloud API.…
★ 335
GenerSpeech
PyTorch Implementation of GenerSpeech (NeurIPS'22): a text-to-speech model towards zero-shot style transfer…
★ 333
Comprehensive-Transformer-TTS
A Non-Autoregressive Transformer based Text-to-Speech, supporting a family of SOTA transformers with…
★ 328
Expressive-FastSpeech2
PyTorch Implementation of Non-autoregressive Expressive (emotional, conversational) TTS based on FastSpeech2,…
★ 322
VocGAN
VocGAN: A High-Fidelity Real-time Vocoder with a Hierarchically-nested Adversarial Network
★ 321
T5Gemma-TTS
Multilingual TTS model with voice cloning and duration control, based on T5Gemma encoder-decoder LLM
★ 311
manim-voiceover
Manim plugin for all things voiceover
★ 308
voiceai
Set of 📝 with 🔗 to help those building Voice AI agents 🎙️🤖
★ 305
axiom-voice-agent
Run a <400ms latency Voice Agent on just 4GB VRAM. Fully offline, no API keys required. Optimized for GTX…
★ 135
Qwen3-TTS-EasyFinetuning
Easy fine-tuning for Qwen3-TTS: Fast voice cloning and high-quality multilingual speech synthesis.
★ 115
Core-AI-Framework-Lab
A practical lab for exploring Apple's Core AI framework, model assets, specialization, and on-device…
★ 62 · GitHub ↗
🔗 Verwandte Familien

Gemessen anhand der von beiden Projekten geteilten GitHub-Themen, gewichtet nach der Seltenheit jedes Themas.