voice-cloning

48 projetos partilham este topic do GitHub

voice-cloning — GPT-SoVITS ★60.2kvoice-cloningReal-Time-Voice-Cloning — ★60.1kTTS — ★45.8kVoxCPM — ★34.4kCosyVoice — ★22.5kVideoLingo — ★17.9kPaddleSpeech — ★12.7kvoice-pro — ★11.3kOmniVoice-Studio — ★9.2kYuE — ★6.3kYouDub-webui — ★5.2kMOSS-TTS — ★3.9kApplio — ★3.5kMARS5-TTS — ★2.8kGenie-TTS — ★1.7kMiniMax-MCP — ★1.5kVibeVoice-ComfyUI — ★1.5kVoice-Cloning-App — ★1.4kopen-speech-corpora — ★1.4kChatterbox-TTS-Server — ★1.4kaudio-webui — ★1.2kTTS-Audio-Suite — ★1.1kIrodori-TTS — ★1.1kStep-Audio-EditX — ★955alexandria-audiobook — ★840CloneTTS — ★751vui — ★737bark-voice-cloning-HuBERT-quantizer — ★710MimikaStudio — ★649chatterbox-tts-api — ★630Pandrator — ★595ComfyUI-VibeVoice — ★588Chatterbox-TTS-Extended — ★573qwen3-tts-apple-silicon — ★537ComfyUI-OmniVoice-TTS — ★520ComfyUI-VoxCPM — ★499mlx-serve — ★380VoxNovel — ★373Cross-Lingual-Voice-Cloning — ★359Dia-TTS-Server — ★352easevoice-trainer — ★351Real-Time-Voice-Cloning★ 60.1kTTS★ 45.8kVoxCPM★ 34.4kCosyVoice★ 22.5kVideoLingo★ 17.9kPaddleSpeech★ 12.7kvoice-pro★ 11.3kOmniVoice-Studio★ 9.2kYuE★ 6.3kYouDub-webui★ 5.2kMOSS-TTS★ 3.9kApplio★ 3.5kMARS5-TTS★ 2.8kGenie-TTS★ 1.7kMiniMax-MCP★ 1.5kVibeVoice-ComfyUI★ 1.5kVoice-Cloning-App★ 1.4kopen-speech-corpora★ 1.4kChatterbox-TTS-Server★ 1.4kaudio-webui★ 1.2kTTS-Audio-Suite★ 1.1kIrodori-TTS★ 1.1kStep-Audio-EditX★ 955alexandria-audiobook★ 840CloneTTS★ 751vui★ 737bark-voice-cloning-HuBER…★ 710MimikaStudio★ 649chatterbox-tts-api★ 630Pandrator★ 595ComfyUI-VibeVoice★ 588Chatterbox-TTS-Extended★ 573qwen3-tts-apple-silicon★ 537ComfyUI-OmniVoice-TTS★ 520ComfyUI-VoxCPM★ 499mlx-serve★ 380VoxNovel★ 373Cross-Lingual-Voice-Clon…★ 359Dia-TTS-Server★ 352easevoice-trainer★ 351

Linhas conectam membros que estão mensuravelmente relacionados entre si. O tamanho do ponto reflete estrelas.

🧬 Membros
GPT-SoVITS
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
★ 60.2k
Real-Time-Voice-Cloning
Clone a voice in 5 seconds to generate arbitrary speech in real-time
★ 60.1k
TTS
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
★ 45.8k
VoxCPM
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life…
★ 34.4k
CosyVoice
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
★ 22.5k
VideoLingo
Netflix-level subtitle cutting, translation, alignment, and even dubbing - one-click fully automated AI video…
★ 17.9k
PaddleSpeech
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation,…
★ 12.7k
voice-pro
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning…
★ 11.3k
OmniVoice-Studio
Local voice clone, video dubbing, dictation and audiobook maker. The open-source ElevenLabs alternative.
★ 9.2k
YuE
YuE: Open Full-song Music Generation Foundation Model, something similar to Suno.ai but open
★ 6.3k
YouDub-webui
★ 5.2k
MOSS-TTS
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS…
★ 3.9k
Applio
A simple, high-quality voice conversion tool focused on ease of use and performance.
★ 3.5k
MARS5-TTS
MARS5 speech model (TTS) from CAMB.AI
★ 2.8k
Genie-TTS
GPT-SoVITS ONNX Inference Engine & Model Converter
★ 1.7k
MiniMax-MCP
Official MiniMax Model Context Protocol (MCP) server that enables interaction with powerful Text to Speech,…
★ 1.5k
VibeVoice-ComfyUI
A comprehensive ComfyUI integration for Microsoft's VibeVoice text-to-speech model, enabling high-quality…
★ 1.5k
Voice-Cloning-App
A Python/Pytorch app for easily synthesising human voices
★ 1.4k
open-speech-corpora
💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
★ 1.4k
Chatterbox-TTS-Server
Self-host the powerful Chatterbox TTS model. This server offers a user-friendly Web UI, flexible API…
★ 1.4k
audio-webui
A webui for different audio related Neural Networks
★ 1.2k
TTS-Audio-Suite
A ComfyUI custom node integration for local multi-engine multi-language Text-to-Speech and Voice Conversion.…
★ 1.1k
Irodori-TTS
A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control
★ 1.1k
Step-Audio-EditX
A powerful 3B-parameter, LLM-based Reinforcement Learning audio edit model excels at editing emotion,…
★ 955
alexandria-audiobook
AI-powered multi-voice audiobook generator — LLM script annotation, voice cloning, voice design, LoRA…
★ 840
CloneTTS
A lightweight, offline Android Text-to-Speech (TTS) engine enabling seamless system-wide voice cloning and…
★ 751
vui
Real-time voice assistant — WebRTC streaming, faster-whisper ASR, local LLM, Vui Nano (300M) TTS. OpenAI…
★ 737
bark-voice-cloning-HuBERT-quantizer
The code for the bark-voicecloning model. Training and inference.
★ 710
MimikaStudio
MimikaStudio - A local-first application for macOS (Apple Silicon) + Agentic MCP Support
★ 649
chatterbox-tts-api
Local, OpenAI-compatible text-to-speech (TTS) API using Chatterbox, enabling users to generate voice cloned…
★ 630
Pandrator
Turn PDFs and EPUBs into audiobooks; subtitles or videos into dubbed videos (including translation), and…
★ 595
ComfyUI-VibeVoice
ComfyUI custom node for the VibeVoice TTS. Expressive, long-form, multi-speaker conversational audio
★ 588
Chatterbox-TTS-Extended
Modified version of Chatterbox that accepts text files as input and no character restrictions. I use it to…
★ 573
qwen3-tts-apple-silicon
Run Qwen3-TTS text-to-speech locally on Mac (M1/M2/M3/M4). Voice cloning, voice design, custom voices. 100%…
★ 537
ComfyUI-OmniVoice-TTS
OmniVoice TTS nodes for ComfyUI - Zero-shot multilingual text-to-speech with voice cloning, voice design, and…
★ 520
ComfyUI-VoxCPM
ComfyUI node for highly expressive speech and realistic zero-shot voice cloning
★ 499
mlx-serve
Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX…
★ 380
VoxNovel
VoxNovel: generate audiobooks giving each character a different voice actor.
★ 373
Cross-Lingual-Voice-Cloning
Tacotron 2 - PyTorch implementation with faster-than-realtime inference modified to enable cross lingual…
★ 359
Dia-TTS-Server
Self-host the powerful Dia TTS model. This server offers a user-friendly Web UI, flexible API endpoints…
★ 352
easevoice-trainer
EaseVoice Trainer is a simple and user-friendly voice cloning and speech model trainer.
★ 351
izwi
Voice AI runtime. Local first transcription, speaker diarization, TTS, and voice cloning with an OpenAI…
★ 351
Vocello
Vocello — a local, private voice studio for Apple Silicon. Write a script, pick or describe a voice, and…
★ 343
ai-avatar-system
🎭 AI Avatar / digital human platform — upload a photo, clone a voice, talk to any face in real time with…
★ 339
UltraEval-Audio
Your faithful, impartial partner for audio evaluation — know yourself, know your rivals.…
★ 311
T5Gemma-TTS
Multilingual TTS model with voice cloning and duration control, based on T5Gemma encoder-decoder LLM
★ 311
ComfyUI-Qwen3-TTS
A ComfyUI custom node suite for Qwen3-TTS, supporting 1.7B and 0.6B models, Custom Voice, Voice Design, Voice…
★ 288
Qwen3-TTS-EasyFinetuning
Easy fine-tuning for Qwen3-TTS: Fast voice cloning and high-quality multilingual speech synthesis.
★ 115
🔗 Familias relacionadas

Medido a partir dos tópicos do GitHub compartilhados por ambos os projetos, ponderado pela raridade de cada tópico.