speech

33 projets partagent ce topic GitHub

speech — MockingBird ★36.9kspeechwhisperX — ★23.1kdatasets — ★21.9kspeech-to-speech — ★13ksilero-vad — ★10.1kannyang — ★6.8kcactus — ★6kwhisper-diarization — ★5.6kstt — ★4.7kultravox — ★4.6kmetavoice-src — ★4.2kMARS5-TTS — ★2.8kIMS-Toucan — ★2.2kopenai-edge-tts — ★2.1kjulius — ★1.9kSALMONN — ★1.5kallosaurus — ★738vits2 — ★642Leaderboard — ★547UniSpeech — ★486huggingsound — ★470docker-whisperX — ★452Ming-UniAudio — ★450speech-recognition-uk — ★439WaveGrad — ★409Freeze-Omni — ★396Stream-Omni — ★392wav2vec2-live — ★378gazelle — ★374MsEdgeTTS — ★335AudioBench — ★319end2end-asr-pytorch — ★304mere-run — ★69whisperX★ 23.1kdatasets★ 21.9kspeech-to-speech★ 13ksilero-vad★ 10.1kannyang★ 6.8kcactus★ 6kwhisper-diarization★ 5.6kstt★ 4.7kultravox★ 4.6kmetavoice-src★ 4.2kMARS5-TTS★ 2.8kIMS-Toucan★ 2.2kopenai-edge-tts★ 2.1kjulius★ 1.9kSALMONN★ 1.5kallosaurus★ 738vits2★ 642Leaderboard★ 547UniSpeech★ 486huggingsound★ 470docker-whisperX★ 452Ming-UniAudio★ 450speech-recognition-uk★ 439WaveGrad★ 409Freeze-Omni★ 396Stream-Omni★ 392wav2vec2-live★ 378gazelle★ 374MsEdgeTTS★ 335AudioBench★ 319end2end-asr-pytorch★ 304mere-run★ 69

Les traits relient les membres réellement apparentés entre eux. La taille des points suit les étoiles.

🧬 Membres
MockingBird
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
★ 36.9k
whisperX
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
★ 23.1k
datasets
🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data…
★ 21.9k
speech-to-speech
Build voice agents with open-source models
★ 13k
silero-vad
Silero VAD: pre-trained enterprise-grade Voice Activity Detector
★ 10.1k
annyang
💬 Speech recognition for your site
★ 6.8k
cactus
Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.
★ 6k
whisper-diarization
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
★ 5.6k
stt
★ 4.7k
ultravox
A fast multimodal LLM for real-time voice
★ 4.6k
metavoice-src
Foundational model for human-like, expressive TTS
★ 4.2k
MARS5-TTS
MARS5 speech model (TTS) from CAMB.AI
★ 2.8k
IMS-Toucan
Controllable and fast Text-to-Speech for over 7000 languages!
★ 2.2k
openai-edge-tts
Free, high-quality text-to-speech API endpoint to replace OpenAI, Azure, or ElevenLabs
★ 2.1k
julius
Open-Source Large Vocabulary Continuous Speech Recognition Engine
★ 1.9k
SALMONN
SALMONN family: A suite of advanced multi-modal LLMs
★ 1.5k
allosaurus
Allosaurus is a pretrained universal phone recognizer for more than 2000 languages
★ 738
vits2
VITS2: Improving Quality and Efficiency of Single-Stage Text-to-Speech with Adversarial Learning and…
★ 642
Leaderboard
SpeechIO Leaderboard: a large, robust, comprehensive, benchmarking platform for Automatic Speech Recognition.
★ 547
UniSpeech
UniSpeech - Large Scale Self-Supervised Learning for Speech
★ 486
huggingsound
HuggingSound: A toolkit for speech-related tasks based on Hugging Face's tools
★ 470
docker-whisperX
Dockerfile for WhisperX: Automatic Speech Recognition with Word-Level Timestamps and Speaker Diarization…
★ 452
Ming-UniAudio
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
★ 450
speech-recognition-uk
🇺🇦 Speech Recognition & Synthesis for Ukrainian
★ 439
WaveGrad
Implementation of WaveGrad high-fidelity vocoder from Google Brain in PyTorch.
★ 409
Freeze-Omni
✨✨Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
★ 396
Stream-Omni
Stream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across…
★ 392
wav2vec2-live
A live speech recognition using Facebooks wav2vec 2.0 model.
★ 378
gazelle
Joint speech-language model - respond directly to audio!
★ 374
MsEdgeTTS
A simple Azure Speech Service module that uses the Microsoft Edge Read Aloud API.…
★ 335
AudioBench
AudioBench: A Universal Benchmark for Audio Large Language Models
★ 319
end2end-asr-pytorch
End-to-End Automatic Speech Recognition on PyTorch
★ 304
mere-run
Run local image, text, speech, vision, music, and video workflows, plus model management and a loopback…
★ 69
🔗 Familles voisines

Mesuré à partir des topics GitHub communs aux deux projets, pondérés par leur rareté.