Home Projects data-preparation
data-preparation
Jupyter Notebook

data-preparation

Code used for sourcing and cleaning the BigScience ROOTS corpus

by bigscience-workshop · GitHub
Stars
Forks
License
Created
Last commit
Category
datasetlarge-language-modelsmultilingualApache-2.0Jupyter Notebook
View on GitHub
In plain words

Source and clean a large dataset for training language models.

You maintain this project?

Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.

Claim this page →
data-preparation — GitHub preview card
📈 Star history
319318
2026-07-202026-08-31
📈 Track data-preparation

Get an email alert on its next release or when it starts trending — never miss the moment.

Free · no card · unsubscribe anytime
Get email alerts →
📄 About

Code used for sourcing and cleaning the BigScience ROOTS corpus

data-preparation has 318 stars on GitHub. It has been forked 43 times. data-preparation is written mainly in Jupyter Notebook. It has been in active development since 2022. data-preparation is available under the Apache-2.0 license. Its main topics are dataset, large-language-models, multilingual.

Frequently asked questions

What is data-preparation?

Code used for sourcing and cleaning the BigScience ROOTS corpus

Is data-preparation open source?

data-preparation is an open-source project. It is released under the Apache-2.0 license.

Is data-preparation free?

Yes. data-preparation is free and open source — you can use, modify and self-host it.

What license does data-preparation use?

data-preparation is available under the Apache-2.0 license.

What language is data-preparation written in?

data-preparation is written mainly in Jupyter Notebook.

🏅 Maintainer of this project?
olud.ai badge — data-preparation

Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.

[![olud.ai](https://olud.ai/badge.php?tool=bigscience-workshop-data-preparation)](https://olud.ai/project/bigscience-workshop-data-preparation.html)
More badge options →
🧬 Shares DNA with🧬 View the DNA map →

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.