synthetic-data

18 projets partagent ce topic GitHub

synthetic-data — machine-learning-for-trading ★20.7ksynthetic-datadata-juicer — ★7kSDV — ★3.6ksynthetic-data-generator — ★2.4kDataDesigner — ★2.2kcurator — ★1.7kdeepfabric — ★880bonito — ★830verbalized-sampling — ★802mostlyai — ★795Copulas — ★652Ditto — ★616EnterpriseRAG-Bench — ★537mtt-distillation — ★442NCFM — ★414be_great — ★366kodcode — ★322SDGym — ★311data-juicer★ 7kSDV★ 3.6ksynthetic-data-generator★ 2.4kDataDesigner★ 2.2kcurator★ 1.7kdeepfabric★ 880bonito★ 830verbalized-sampling★ 802mostlyai★ 795Copulas★ 652Ditto★ 616EnterpriseRAG-Bench★ 537mtt-distillation★ 442NCFM★ 414be_great★ 366kodcode★ 322SDGym★ 311

Les traits relient les membres réellement apparentés entre eux. La taille des points suit les étoiles.

🧬 Membres
machine-learning-for-trading
Code for Machine Learning for Trading, 3rd edition — from data sourcing to live execution.
★ 20.7k
data-juicer
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
★ 7k
SDV
Synthetic data generation for tabular data
★ 3.6k
synthetic-data-generator
SDG is a specialized framework designed to generate high-quality structured tabular data.
★ 2.4k
DataDesigner
🎨 NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data.
★ 2.2k
curator
Synthetic data curation for post-training and structured data extraction
★ 1.7k
deepfabric
Generate High-Quality Synthetics, Train, Measure, and Evaluate in a Single Pipeline
★ 880
bonito
A lightweight library for generating synthetic instruction tuning datasets for your data without GPT.
★ 830
verbalized-sampling
Verbalized Sampling, a training-free prompting strategy to mitigate mode collapse in LLMs by requesting…
★ 802
mostlyai
Synthetic Data SDK ✨
★ 795
Copulas
A library to model multivariate data using copulas.
★ 652
Ditto
[CVPR'26 Highlight] Ditto: Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset
★ 616
EnterpriseRAG-Bench
Dataset and benchmark for RAG on company internal documents.
★ 537
mtt-distillation
Official code for our CVPR '22 paper "Dataset Distillation by Matching Training Trajectories"
★ 442
NCFM
Official PyTorch implementation of the paper "Dataset Distillation with Neural Characteristic Function: A…
★ 414
be_great
A novel approach for synthesizing tabular data using pretrained large language models
★ 366
kodcode
✨ A synthetic dataset generation framework that produces diverse coding questions and verifiable solutions…
★ 322
SDGym
Benchmarking synthetic data generation methods.
★ 311
🔗 Familles voisines

Mesuré à partir des topics GitHub communs aux deux projets, pondérés par leur rareté.