Home Projects data-juicer
data-juicer
Python

data-juicer

Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷

by datajuicer · GitHub
Stars
Forks
License
Created
Last commit
Category
Language
datadata-analysisdata-pipelineApache-2.0Python
View on GitHub
In plain words

Transform raw data into organized, AI-ready information with a flexible data processing tool that scales easily.

You maintain this project?

Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.

Claim this page →
data-juicer — GitHub preview card
📈 Star history
7.0k6.6k
2026-07-042026-08-31
📈 Track data-juicer

Get an email alert on its next release or when it starts trending — never miss the moment.

Free · no card · unsubscribe anytime
Get email alerts →
📄 About

Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷

data-juicer has 7k stars on GitHub. It has been forked 412 times. data-juicer is written mainly in Python. It has been in active development since 2023. data-juicer is available under the Apache-2.0 license. Its main topics are data, data-analysis, data-pipeline, data-processing.

Frequently asked questions

What is data-juicer?

Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷

Is data-juicer open source?

data-juicer is an open-source project. It is released under the Apache-2.0 license.

Is data-juicer free?

Yes. data-juicer is free and open source — you can use, modify and self-host it.

What license does data-juicer use?

data-juicer is available under the Apache-2.0 license.

What language is data-juicer written in?

data-juicer is written mainly in Python.

🏅 Maintainer of this project?
olud.ai badge — data-juicer

Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.

[![olud.ai](https://olud.ai/badge.php?tool=datajuicer-data-juicer)](https://olud.ai/project/datajuicer-data-juicer.html)
More badge options →
🧬 Shares DNA with🧬 View the DNA map →
Python
This repository helps you learn Python and Machine Learning from scratch.
2.2k · data
sharesdata-visualizationdata
streamlit
Streamlit — A faster way to build and share data apps.
45.3k · data-analysis
sharesdata-visualizationdata-analysis
gradio
Build and share delightful machine learning apps, all in Python. 🌟 Star to support our work!
43.2k · data-analysis
sharesdata-visualizationdata-analysis
machine_learning_complete
A comprehensive machine learning repository containing 30+ notebooks on different concepts, alg…
5k · computer-vision
sharesdata-visualizationdata-analysis
mito
Jupyter extensions that help you write code faster: Context aware AI Chat, Autocomplete, and Sp…
2.6k · ai
sharesdatadata-analysis
Data-Science-Hacks
Data Science Hacks consists of tips, tricks to help you become a better data scientist. Data sc…
432 · computer-vision
sharesdatadata-analysis
deepnote
Deepnote is a drop-in replacement for Jupyter with an AI-first design, sleek UI, new blocks, an…
3k · artificial-intelligence
sharesdatadata-analysis
Book6_First-Course-in-Data-Science
Book_6_《数据有道》 | 鸢尾花书:从加减乘除到机器学习;欢迎大家批评指正!纠错多的同学会得到赠书感谢!
2.7k · data
sharesdata-visualizationdata
ml-workspace
🛠 All-in-one web-based IDE specialized for machine learning and data science.
3.5k · anaconda
sharesdata-visualizationdata-analysis
responsible-ai-toolbox
Responsible AI Toolbox is a suite of tools providing model and data exploration and assessment…
1.8k · data-analysis
sharesdata-visualizationdata-analysis

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.