Home Projects doremi
doremi
HTML

doremi

Pytorch implementation of DoReMi, a method for optimizing the data mixture weights in language modeling datasets

by sangmichaelxie · GitHub
Stars
Forks
Trending
License
Created
Last commit
Language
data-centric-machine-learninglarge-language-modelsnlpMITHTML
View on GitHub
In plain words

Optimize data mixtures for training language models using a PyTorch implementation of the DoReMi algorithm.

You maintain this project?

Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.

Claim this page →
doremi — GitHub preview card
📈 Star history
358357
2026-07-202026-08-31
📈 Track doremi

Get an email alert on its next release or when it starts trending — never miss the moment.

Free · no card · unsubscribe anytime
Get email alerts →
📄 About

Pytorch implementation of DoReMi, a method for optimizing the data mixture weights in language modeling datasets

doremi has 358 stars on GitHub. It has been forked 35 times. doremi is written mainly in HTML. It has been in active development since 2023. doremi is available under the MIT license. Its main topics are data-centric-machine-learning, large-language-models, nlp.

Frequently asked questions

What is doremi?

Pytorch implementation of DoReMi, a method for optimizing the data mixture weights in language modeling datasets

Is doremi open source?

doremi is an open-source project. It is released under the MIT license.

Is doremi free?

Yes. doremi is free and open source — you can use, modify and self-host it.

What license does doremi use?

doremi is available under the MIT license.

What language is doremi written in?

doremi is written mainly in HTML.

🏅 Maintainer of this project?
olud.ai badge — doremi

Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.

[![olud.ai](https://olud.ai/badge.php?tool=sangmichaelxie-doremi)](https://olud.ai/project/sangmichaelxie-doremi.html)
More badge options →
🧬 Shares DNA with

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.