Home Projects data-prep-kit
data-prep-kit
HTML

data-prep-kit

Open source project for data preparation for GenAI applications

by data-prep-kit · GitHub
Stars
Forks
Trending
License
Created
Last commit
Category
Language
code-qualitydatadata-prepApache-2.0HTML
View on GitHub
In plain words

Prepare and clean unstructured data for training AI models, scalable from personal computers to data centers.

You maintain this project?

Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.

Claim this page →
data-prep-kit — GitHub preview card
📈 Star history
956948
2026-07-202026-08-31
📈 Track data-prep-kit

Get an email alert on its next release or when it starts trending — never miss the moment.

Free · no card · unsubscribe anytime
Get email alerts →
📄 About

Open source project for data preparation for GenAI applications

data-prep-kit has 956 stars on GitHub. It has been forked 253 times. data-prep-kit is written mainly in HTML. It has been in active development since 2024. data-prep-kit is available under the Apache-2.0 license. Its main topics are code-quality, data, data-prep, data-preparation.

Frequently asked questions

What is data-prep-kit?

Open source project for data preparation for GenAI applications

Is data-prep-kit open source?

data-prep-kit is an open-source project. It is released under the Apache-2.0 license.

Is data-prep-kit free?

Yes. data-prep-kit is free and open source — you can use, modify and self-host it.

What license does data-prep-kit use?

data-prep-kit is available under the Apache-2.0 license.

What language is data-prep-kit written in?

data-prep-kit is written mainly in HTML.

🏅 Maintainer of this project?
olud.ai badge — data-prep-kit

Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.

[![olud.ai](https://olud.ai/badge.php?tool=data-prep-kit-data-prep-kit)](https://olud.ai/project/data-prep-kit-data-prep-kit.html)
More badge options →
🧬 Shares DNA with🧬 View the DNA map →

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.