Home Projects SpeechT5
SpeechT5
Python

SpeechT5

Unified-Modal Speech-Text Pre-Training for Spoken Language Processing

by microsoft · GitHub
Stars
Forks
License
Created
Last commit
Language
speech-pretrainingspeech-recognitionspeech-synthesisMITPython
View on GitHub
In plain words

Train models to understand and process spoken language, converting speech to text and vice versa.

You maintain this project?

Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.

Claim this page →
SpeechT5 — GitHub preview card
📈 Star history
1 4481 447
2026-07-202026-08-31
📈 Track SpeechT5

Get an email alert on its next release or when it starts trending — never miss the moment.

Free · no card · unsubscribe anytime
Get email alerts →
📄 About

Unified-Modal Speech-Text Pre-Training for Spoken Language Processing

SpeechT5 has 1.4k stars on GitHub. It has been forked 134 times. SpeechT5 is written mainly in Python. It has been in active development since 2022. SpeechT5 is available under the MIT license. Its main topics are speech-pretraining, speech-recognition, speech-synthesis, speech-text-pretraining.

Frequently asked questions

What is SpeechT5?

Unified-Modal Speech-Text Pre-Training for Spoken Language Processing

Is SpeechT5 open source?

SpeechT5 is an open-source project. It is released under the MIT license.

Is SpeechT5 free?

Yes. SpeechT5 is free and open source — you can use, modify and self-host it.

What license does SpeechT5 use?

SpeechT5 is available under the MIT license.

What language is SpeechT5 written in?

SpeechT5 is written mainly in Python.

🏅 Maintainer of this project?
olud.ai badge — SpeechT5

Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.

[![olud.ai](https://olud.ai/badge.php?tool=microsoft-speecht5)](https://olud.ai/project/microsoft-speecht5.html)
More badge options →
🧬 Shares DNA with🧬 View the DNA map →
local-talking-llm
A talking LLM that runs on your own computer without needing the internet.
878 · chatbot
sharesspeech-synthesisspeech-recognition
Speech-Backbones
This is the main repository of open-sourced speech technology by Huawei Noah's Ark Lab.
604 · speech-processing
sharesspeech-synthesisspeech-recognition
Irene-Voice-Assistant
Ирина - русский голосовой ассистент для работы оффлайн. Поддерживает скиллы через плагины.
1.1k · python
sharesspeech-synthesisspeech-recognition
speech-recognition-uk
🇺🇦 Speech Recognition & Synthesis for Ukrainian
439 · speech
sharesspeech-synthesisspeech-recognition
artyom.js
A voice control - voice commands - speech recognition and speech synthesis javascript library.…
1.3k · recognition
sharesspeech-synthesisspeech-recognition
Ming-UniAudio
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Represen…
450 · speech
sharesspeech-synthesisspeech-recognition
Freeze-Omni
✨✨Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
395 · large-language-models
sharesspeech-synthesisspeech-recognition
kaldi-gstreamer-server
Real-time full-duplex speech recognition server, based on the Kaldi toolkit and the GStreamer f…
1.1k · speech-recognition
sharesspeech-recognition
FastASR
这是一个用C++实现ASR推理的项目,它依赖很少,安装也很简单,推理速度很快,在树莓派4B等ARM平台也可以流畅的运行。 支持的模型是由Google的Transformer模型中优化而来,数…
553 · speech-recognition
sharesspeech-recognition
dragonfly
Speech recognition framework allowing powerful Python-based scripting and extension of Dragon N…
413 · python
sharesspeech-recognition

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.