Home Projects tokenizers
tokenizers
Rust

tokenizers

💥 Fast State-of-the-Art Tokenizers optimized for Research and Production

by huggingface · GitHub
Stars
Forks
License
Created
Last commit
Category
Language
bertgptlanguage-modelApache-2.0Rust
View on GitHub
In plain words

Train and use advanced text processing tools quickly and efficiently for research and production.

You maintain this project?

Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.

Claim this page →
tokenizers — GitHub preview card
📈 Star history
10.89k10.86k
2026-07-042026-08-31
📈 Track tokenizers

Get an email alert on its next release or when it starts trending — never miss the moment.

Free · no card · unsubscribe anytime
Get email alerts →
📄 About

💥 Fast State-of-the-Art Tokenizers optimized for Research and Production

tokenizers has 10.9k stars on GitHub. It has been forked 1.1k times. tokenizers is written mainly in Rust. It has been in active development since 2019. tokenizers is available under the Apache-2.0 license. Its main topics are bert, gpt, language-model, natural-language-processing.

Frequently asked questions

What is tokenizers?

💥 Fast State-of-the-Art Tokenizers optimized for Research and Production

Is tokenizers open source?

tokenizers is an open-source project. It is released under the Apache-2.0 license.

Is tokenizers free?

Yes. tokenizers is free and open source — you can use, modify and self-host it.

What license does tokenizers use?

tokenizers is available under the Apache-2.0 license.

What language is tokenizers written in?

tokenizers is written mainly in Rust.

🏅 Maintainer of this project?
olud.ai badge — tokenizers

Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.

[![olud.ai](https://olud.ai/badge.php?tool=huggingface-tokenizers)](https://olud.ai/project/huggingface-tokenizers.html)
More badge options →
🧬 Shares DNA with🧬 View the DNA map →

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.