Pytorch implementation of DoReMi, a method for optimizing the data mixture weights in language modeling datasets
Optimize data mixtures for training language models using a PyTorch implementation of the DoReMi algorithm.
Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.
Get an email alert on its next release or when it starts trending — never miss the moment.
Free · no card · unsubscribe anytimePytorch implementation of DoReMi, a method for optimizing the data mixture weights in language modeling datasets
doremi has 358 stars on GitHub. It has been forked 35 times. doremi is written mainly in HTML. It has been in active development since 2023. doremi is available under the MIT license. Its main topics are data-centric-machine-learning, large-language-models, nlp.
Pytorch implementation of DoReMi, a method for optimizing the data mixture weights in language modeling datasets
doremi is an open-source project. It is released under the MIT license.
Yes. doremi is free and open source — you can use, modify and self-host it.
doremi is available under the MIT license.
doremi is written mainly in HTML.
Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.
[](https://olud.ai/project/sangmichaelxie-doremi.html)
Measured from GitHub topics shared by both projects, weighted by how rare each topic is.