Home Projects bonito
bonito
Python

bonito

A lightweight library for generating synthetic instruction tuning datasets for your data without GPT.

by BatsResearch · GitHub
Stars
Forks
Created
Last commit
Language
domain-adaptationgptllmBSD-3-ClausePython
View on GitHub
In plain words

Generate training datasets from unannotated text for AI models easily.

You maintain this project?

Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.

Claim this page →
bonito — GitHub preview card
📈 Star history
831830
2026-07-202026-08-31
📈 Track bonito

Get an email alert on its next release or when it starts trending — never miss the moment.

Free · no card · unsubscribe anytime
Get email alerts →
📄 About

A lightweight library for generating synthetic instruction tuning datasets for your data without GPT.

bonito has 830 stars on GitHub. It has been forked 56 times. bonito is written mainly in Python. It has been in active development since 2024. bonito is available under the BSD-3-Clause license. Its main topics are domain-adaptation, gpt, llm, synthetic-data.

Frequently asked questions

What is bonito?

A lightweight library for generating synthetic instruction tuning datasets for your data without GPT.

Is bonito open source?

bonito is an open-source project. It is released under the BSD-3-Clause license.

Is bonito free?

Yes. bonito is free and open source — you can use, modify and self-host it.

What license does bonito use?

bonito is available under the BSD-3-Clause license.

What language is bonito written in?

bonito is written mainly in Python.

🏅 Maintainer of this project?
olud.ai badge — bonito

Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.

[![olud.ai](https://olud.ai/badge.php?tool=batsresearch-bonito)](https://olud.ai/project/batsresearch-bonito.html)
More badge options →
🧬 Shares DNA with🧬 View the DNA map →
pytorch-adapt
Domain adaptation made easy. Fully featured, modular, and customizable.
397 · computer-vision
sharesdomain-adaptation
pykale
Knowledge-Aware machine LEarning (KALE): accessible machine learning from multiple sources for…
486 · computer-vision
sharesdomain-adaptation
Exclusively-Dark-Image-Dataset
Exclusively Dark (ExDARK) dataset which to the best of our knowledge, is the largest collection…
635 · computer-vision
sharesdomain-adaptation
mtt-distillation
Official code for our CVPR '22 paper "Dataset Distillation by Matching Training Trajectories"
442 · artificial-intelligence
sharessynthetic-data
powerful-benchmarker
A library for ML benchmarking. It's powerful.
441 · benchmarking
sharesdomain-adaptation
Ditto
[CVPR'26 Highlight] Ditto: Scaling Instruction-Based Video Editing with a High-Quality Syntheti…
616 · diffusion-models
sharessynthetic-data
AdaptSegNet
Learning to Adapt Structured Output Space for Semantic Segmentation, CVPR 2018 (spotlight)
859 · adversarial-learning
sharesdomain-adaptation
gpl
Powerful unsupervised domain adaptation method for dense retrieval. Requires only unlabeled cor…
343 · bert
sharesdomain-adaptation
transferlearning
Transfer learning / domain adaptation / domain generalization / multi-task learning etc. Papers…
14.3k · deep-learning
sharesdomain-adaptation
Transfer-Learning-Library
Transfer Learning Library for Domain Adaptation, Task Adaptation, and Domain Generalization
3.9k · adversarial-learning
sharesdomain-adaptation

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.