speech

86 proyectos comparten este topic de GitHub

speech — TTS ★45.8kspeechMockingBird — ★36.9kVoxCPM — ★34.4kwhisperX — ★23.3kdatasets — ★21.8kkaldi — ★15.4kAudioGPT — ★10.2kTTS — ★10.2kmodelscope — ★9.1kEmotiVoice — ★8.5kspeech-to-speech — ★7kmodels — ★6.9kannyang — ★6.8ksilero-models — ★6kwhisper-diarization — ★5.6kcactus — ★5.5kstt — ★4.7kDeepFilterNet — ★4.5kultravox — ★4.5kClearerVoice-Studio — ★4.3kmetavoice-src — ★4.2kAmazing-Python-Scripts — ★3.6kApplio — ★3.5kwhisper-asr-webservice — ★3.3kaudio — ★2.9klingvo — ★2.9kaeneas — ★2.9kwhisper-timestamped — ★2.8kMARS5-TTS — ★2.8kspeechgpt — ★2.8kgTTS — ★2.6kpytorch-kaldi — ★2.4kIMS-Toucan — ★2.2kopenai-edge-tts — ★2kjulius — ★1.9kSALMONN — ★1.5kSpeech-Emotion-Analyzer — ★1.4kStreamSpeech — ★1.3klhotse — ★1.1kconformer — ★1.1kpykaldi — ★1kMockingBird★ 36.9kVoxCPM★ 34.4kwhisperX★ 23.3kdatasets★ 21.8kkaldi★ 15.4kAudioGPT★ 10.2kTTS★ 10.2kmodelscope★ 9.1kEmotiVoice★ 8.5kspeech-to-speech★ 7kmodels★ 6.9kannyang★ 6.8ksilero-models★ 6kwhisper-diarization★ 5.6kcactus★ 5.5kstt★ 4.7kDeepFilterNet★ 4.5kultravox★ 4.5kClearerVoice-Studio★ 4.3kmetavoice-src★ 4.2kAmazing-Python-Scripts★ 3.6kApplio★ 3.5kwhisper-asr-webservice★ 3.3kaudio★ 2.9klingvo★ 2.9kaeneas★ 2.9kwhisper-timestamped★ 2.8kMARS5-TTS★ 2.8kspeechgpt★ 2.8kgTTS★ 2.6kpytorch-kaldi★ 2.4kIMS-Toucan★ 2.2kopenai-edge-tts★ 2kjulius★ 1.9kSALMONN★ 1.5kSpeech-Emotion-Analyzer★ 1.4kStreamSpeech★ 1.3klhotse★ 1.1kconformer★ 1.1kpykaldi★ 1k

Las líneas conectan a los miembros que están mediblemente relacionados entre sí. El tamaño de los puntos refleja las estrellas.

🧬 Miembros
TTS
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
★ 45.8k
MockingBird
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
★ 36.9k
VoxCPM
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life…
★ 34.4k
whisperX
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
★ 23.3k
datasets
🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data…
★ 21.8k
kaldi
kaldi-asr/kaldi is the official location of the Kaldi project.
★ 15.4k
AudioGPT
AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
★ 10.2k
TTS
:robot: :speech_balloon: Deep learning for Text to Speech (Discussion forum:…
★ 10.2k
modelscope
ModelScope: bring the notion of Model-as-a-Service to life.
★ 9.1k
EmotiVoice
EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine
★ 8.5k
speech-to-speech
Build local voice agents with open-source models
★ 7k
models
Officially maintained, supported by PaddlePaddle, including CV, NLP, Speech, Rec, TS, big models and so on.
★ 6.9k
annyang
💬 Speech recognition for your site
★ 6.8k
silero-models
Silero Models: pre-trained text-to-speech models made embarrassingly simple
★ 6k
whisper-diarization
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
★ 5.6k
cactus
Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.
★ 5.5k
stt
★ 4.7k
DeepFilterNet
Noise supression using deep filtering
★ 4.5k
ultravox
A fast multimodal LLM for real-time voice
★ 4.5k
ClearerVoice-Studio
An AI-Powered Speech Processing Toolkit and Open Source SOTA Pretrained Models, Supporting Speech…
★ 4.3k
metavoice-src
Foundational model for human-like, expressive TTS
★ 4.2k
Amazing-Python-Scripts
🚀 Curated collection of Amazing Python scripts from Basics to Advance with automation task scripts.
★ 3.6k
Applio
A simple, high-quality voice conversion tool focused on ease of use and performance.
★ 3.5k
whisper-asr-webservice
OpenAI Whisper ASR Webservice API
★ 3.3k
audio
Data manipulation and transformation for audio signal processing, powered by PyTorch
★ 2.9k
lingvo
Lingvo
★ 2.9k
aeneas
aeneas is a Python/C library and a set of tools to automagically synchronize audio and text (aka forced…
★ 2.9k
whisper-timestamped
Multilingual Automatic Speech Recognition with word-level timestamps and confidence
★ 2.8k
MARS5-TTS
MARS5 speech model (TTS) from CAMB.AI
★ 2.8k
speechgpt
💬 SpeechGPT is a web application that enables you to converse with ChatGPT.
★ 2.8k
gTTS
Python library and CLI tool to interface with Google Translate's text-to-speech API
★ 2.6k
pytorch-kaldi
pytorch-kaldi is a project for developing state-of-the-art DNN/RNN hybrid speech recognition systems. The DNN…
★ 2.4k
IMS-Toucan
Controllable and fast Text-to-Speech for over 7000 languages!
★ 2.2k
openai-edge-tts
Free, high-quality text-to-speech API endpoint to replace OpenAI, Azure, or ElevenLabs
★ 2k
julius
Open-Source Large Vocabulary Continuous Speech Recognition Engine
★ 1.9k
SALMONN
SALMONN family: A suite of advanced multi-modal LLMs
★ 1.5k
Speech-Emotion-Analyzer
The neural network model is capable of detecting five different male/female emotions from audio speeches.…
★ 1.4k
StreamSpeech
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech…
★ 1.3k
lhotse
Tools for handling multimodal data in machine learning projects.
★ 1.1k
conformer
[Unofficial] PyTorch implementation of "Conformer: Convolution-augmented Transformer for Speech Recognition"…
★ 1.1k
pykaldi
A Python wrapper for Kaldi
★ 1k
CrisperWhisper
Verbatim Automatic Speech Recognition with improved word-level timestamps and filler detection
★ 1k
FaceFormer
[CVPR 2022] FaceFormer: Speech-Driven 3D Facial Animation with Transformers
★ 915
FireRedTTS
An Open-Sourced LLM-empowered Foundation TTS System
★ 909
diffwave
DiffWave is a fast, high-quality neural vocoder and waveform synthesizer.
★ 885
PPASR
★ 871
VAD
Voice activity detection (VAD) toolkit including DNN, bDNN, LSTM and ACAM based VAD. We also provide our…
★ 869
cboard
Augmentative and Alternative Communication (AAC) system with text-to-speech for the browser
★ 743
allosaurus
Allosaurus is a pretrained universal phone recognizer for more than 2000 languages
★ 737
MASR
★ 728
openspeech
Open-Source Toolkit for End-to-End Speech Recognition leveraging PyTorch-Lightning and Hydra.
★ 716
BabelDuck
Beginner-friendly AI conversation practice application
★ 691
SpecAugment
A Implementation of SpecAugment with Tensorflow & Pytorch, introduced by Google Brain
★ 655
vits2
VITS2: Improving Quality and Efficiency of Single-Stage Text-to-Speech with Adversarial Learning and…
★ 643
sonus
:speech_balloon: /so.nus/ STT (speech to text) for Node with offline hotword detection
★ 638
chatterbox-tts-api
Local, OpenAI-compatible text-to-speech (TTS) API using Chatterbox, enabling users to generate voice cloned…
★ 630
neural_sp
End-to-end ASR/LM implementation with PyTorch
★ 594
Leaderboard
SpeechIO Leaderboard: a large, robust, comprehensive, benchmarking platform for Automatic Speech Recognition.
★ 547
java-speech-api
The J.A.R.V.I.S. Speech API is designed to be simple and efficient, using the speech engines created by…
★ 542
awesome-kaldi
This is a list of features, scripts, blogs and resources for better using Kaldi ( http://kaldi-asr.org/ )
★ 536
Awesome-Singing-Voice-Synthesis-and-Singing-Voice-Conversion
A paper and project list about the cutting edge Speech Synthesis, Text-to-Speech (TTS), Singing Voice…
★ 487
UniSpeech
UniSpeech - Large Scale Self-Supervised Learning for Speech
★ 486
huggingsound
HuggingSound: A toolkit for speech-related tasks based on Hugging Face's tools
★ 469
speech_dataset
The dataset of Speech Recognition
★ 464
docker-whisperX
Dockerfile for WhisperX: Automatic Speech Recognition with Word-Level Timestamps and Speaker Diarization…
★ 454
Ming-UniAudio
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
★ 451
echogarden
Cross-platform speech toolset, used from the command-line or as a Node.js library. Includes a variety of…
★ 444
speech-recognition-uk
🇺🇦 Speech Recognition & Synthesis for Ukrainian
★ 440
tevr-asr-tool
State-of-the-art (ranked #1 Aug 2022) German Speech Recognition in 284 lines of C++. This is a 100% private…
★ 410
WaveGrad
Implementation of WaveGrad high-fidelity vocoder from Google Brain in PyTorch.
★ 409
potato
potato: the portable annotation tool
★ 399
Stream-Omni
Stream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across…
★ 390
Freeze-Omni
✨✨Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
★ 388
NanoLLM
Optimized local inference for LLMs with HuggingFace-like APIs for quantization, vision/language models,…
★ 381
wav2vec2-live
A live speech recognition using Facebooks wav2vec 2.0 model.
★ 379
parakeet-rs
very fast speech-to-text, diarization, streaming (even in CPU) with NVIDIA Parakeet in Rust
★ 378
speechbrain.github.io
The SpeechBrain project aims to build a novel speech toolkit fully based on PyTorch. With SpeechBrain users…
★ 374
gazelle
Joint speech-language model - respond directly to audio!
★ 374
voice_activity_detection
Voice Activity Detection based on Deep Learning & TensorFlow
★ 373
AIUI
AIUI is a platform enabling seamless two-way verbal communication with AI.
★ 356
LangHelper
Striving to create a great Application with full functions of learning languages by ChatGPT, TTS, STT and…
★ 349
MsEdgeTTS
A simple Azure Speech Service module that uses the Microsoft Edge Read Aloud API.…
★ 335
AudioBench
AudioBench: A Universal Benchmark for Audio Large Language Models
★ 319
obsidian-edge-tts
Free, high quality text-to-speech for your Obsidian notes, leveraging Microsoft Edge's Read Aloud API.
★ 310
xcodec
AAAI 2025: Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model
★ 308
end2end-asr-pytorch
End-to-End Automatic Speech Recognition on PyTorch
★ 304
🔗 Familias relacionadas

Medido a partir de los temas de GitHub compartidos por ambos proyectos, ponderado por cuán raros son cada uno de los temas.