A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
Develop speech recognition and text-to-speech applications using a flexible AI framework.
git clone https://github.com/NVIDIA-NeMo/Speech.git cd Speech uv sync --extra all --extra cu13 # CUDA 13.x (recommended) — use --extra cu12 for CUDA 12.x
Excerpts from the project README on GitHub. Copyright and licensing remain with the respective authors.
Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.
Get an email alert on its next release or when it starts trending — never miss the moment.
Free · no card · unsubscribe anytimeA scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
Speech has 18.4k stars on GitHub. It has been forked 3.6k times. Speech is written mainly in Python. It has been in active development since 2019. Speech is available under the Apache-2.0 license. Its main topics are asr, deeplearning, generative-ai, machine-translation.
Compare with the paid tools it replacesA scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
Speech is an open-source project. It is released under the Apache-2.0 license.
Yes. Speech is free and open source — you can use, modify and self-host it.
Speech is available under the Apache-2.0 license.
Speech is written mainly in Python.
Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.
[](https://olud.ai/project/nvidia-nemo-speech.html)
Measured from GitHub topics shared by both projects, weighted by how rare each topic is.