data

53 progetti condividono questo topic GitHub

data — Scrapling ★77.6kdatallama_index — ★51.9kmetabase — ★49kpandas-ai — ★23.8kairbyte — ★22kelectric — ★10.3kmage-ai — ★8.8kmachine-learning-roadmap — ★7.9kDataFlow — ★7.8kdata-juicer — ★7kmachine-learning-mindmap — ★6.3kEdit-Banana — ★5.5ksuperduper — ★5.3kllm-datasets — ★4.8kdatasets — ★4.6kDeepAnalyze — ★4.6kpreswald — ★4.3kdocetl — ★4.1kLazyLLM — ★3.9kAnyCrawl — ★3.4kspiceai — ★3kstats — ★3kweld — ★3kdeepnote — ★3kquant-mind — ★2.8kdatasets — ★2.8kBook6_First-Course-in-Data-Science — ★2.7kmito — ★2.6kDeepBI — ★2.4ksketch — ★2.3kawesome-streamlit — ★2.3kTigerBot — ★2.3kPython — ★2.2kfree-ai-resources — ★1.9kdiffgram — ★1.9kCurator — ★1.7kdurable-streams — ★1.7klotus — ★1.7ksynthetic-data-kit — ★1.6kcovid19_scenarios — ★1.4klhotse — ★1.1kllama_index★ 51.9kmetabase★ 49kpandas-ai★ 23.8kairbyte★ 22kelectric★ 10.3kmage-ai★ 8.8kmachine-learning-roadmap★ 7.9kDataFlow★ 7.8kdata-juicer★ 7kmachine-learning-mindmap★ 6.3kEdit-Banana★ 5.5ksuperduper★ 5.3kllm-datasets★ 4.8kdatasets★ 4.6kDeepAnalyze★ 4.6kpreswald★ 4.3kdocetl★ 4.1kLazyLLM★ 3.9kAnyCrawl★ 3.4kspiceai★ 3kstats★ 3kweld★ 3kdeepnote★ 3kquant-mind★ 2.8kdatasets★ 2.8kBook6_First-Course-in-Da…★ 2.7kmito★ 2.6kDeepBI★ 2.4ksketch★ 2.3kawesome-streamlit★ 2.3kTigerBot★ 2.3kPython★ 2.2kfree-ai-resources★ 1.9kdiffgram★ 1.9kCurator★ 1.7kdurable-streams★ 1.7klotus★ 1.7ksynthetic-data-kit★ 1.6kcovid19_scenarios★ 1.4klhotse★ 1.1k

Le linee collegano membri che sono misurabilmente correlati tra loro. La dimensione dei punti riflette le stelle.

🧬 Membri
Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale…
★ 77.6k
llama_index
LlamaIndex is the leading document agent and OCR platform
★ 51.9k
metabase
The easy-to-use open source Business Intelligence and Embedded Analytics tool that lets everyone work with…
★ 49k
pandas-ai
Chat with your database or your datalake (SQL, CSV, parquet). PandasAI makes data analysis conversational…
★ 23.8k
airbyte
Open-source data movement for ELT pipelines and AI agents — from APIs, databases & files to warehouses,…
★ 22k
electric
The agent platform built on sync.
★ 10.3k
mage-ai
🧙 Build, run, and manage data pipelines for integrating and transforming data.
★ 8.8k
machine-learning-roadmap
A roadmap connecting many of the most important concepts in machine learning, how to learn them and what…
★ 7.9k
DataFlow
Easy Data Preparation with latest LLMs-based Operators and Pipelines.
★ 7.8k
data-juicer
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
★ 7k
machine-learning-mindmap
A mindmap summarising Machine Learning concepts, from Data Analysis to Deep Learning.
★ 6.3k
Edit-Banana
Edit Banana: A framework for converting statistical formats into editable.
★ 5.5k
superduper
Superduper: End-to-end framework for building custom AI applications and agents.
★ 5.3k
llm-datasets
Curated list of datasets and tools for post-training.
★ 4.8k
datasets
TFDS is a collection of datasets ready to use with TensorFlow, Jax, ...
★ 4.6k
DeepAnalyze
★ 4.6k
preswald
Preswald is a WASM packager for Python-based interactive data apps: bundle full complex data workflows,…
★ 4.3k
docetl
A system for agentic LLM-powered data processing and ETL
★ 4.1k
LazyLLM
Easiest and laziest way for building multi-agent LLMs applications.
★ 3.9k
AnyCrawl
AnyCrawl 🚀: A Node.js/TypeScript crawler that turns websites into LLM-ready data and extracts structured…
★ 3.4k
spiceai
Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query,…
★ 3k
stats
A well tested and comprehensive Golang statistics library package with no dependencies.
★ 3k
weld
High-performance runtime for data analytics applications
★ 3k
deepnote
Deepnote is a drop-in replacement for Jupyter with an AI-first design, sleek UI, new blocks, and native data…
★ 3k
quant-mind
QuantMind is an agent-native knowledge extraction and retrieval framework for quantitative finance.
★ 2.8k
datasets
🎁 7,400,000+ Unsplash images made available for research and machine learning
★ 2.8k
Book6_First-Course-in-Data-Science
★ 2.7k
mito
Jupyter extensions that help you write code faster: Context aware AI Chat, Autocomplete, and Spreadsheet
★ 2.6k
DeepBI
LLM based data scientist, AI native data application. AI-driven infinite thinking redefines BI.
★ 2.4k
sketch
AI code-writing assistant that understands data content
★ 2.3k
awesome-streamlit
The purpose of this project is to share knowledge on how awesome Streamlit is and can be
★ 2.3k
TigerBot
TigerBot: A multi-language multi-task LLM
★ 2.3k
Python
This repository helps you learn Python and Machine Learning from scratch.
★ 2.2k
free-ai-resources
🚀 FREE AI Resources - 🎓 Courses, 👷 Jobs, 📝 Blogs, 🔬 AI Research, and many more - for everyone!
★ 1.9k
diffgram
The AI Datastore for Schemas, BLOBs, and Predictions. Use with your apps or integrate built-in Human…
★ 1.9k
Curator
Scalable data pre processing and curation toolkit for LLMs
★ 1.7k
durable-streams
The data primitive for the agent loop.
★ 1.7k
lotus
Optimized Agentic and LLM Bulk Processing Over Your Data
★ 1.7k
synthetic-data-kit
Tool for generating high quality Synthetic datasets
★ 1.6k
covid19_scenarios
Models of COVID-19 outbreak trajectories and hospital demand
★ 1.4k
lhotse
Tools for handling multimodal data in machine learning projects.
★ 1.1k
mcap
MCAP is a modular, performant, and serialization-agnostic container file format, useful for pub/sub and…
★ 1k
data-prep-kit
Open source project for data preparation for GenAI applications
★ 956
vectordb
Epsilla is a high performance Vector Database Management System
★ 875
VAD
Voice activity detection (VAD) toolkit including DNN, bDNN, LSTM and ACAM based VAD. We also provide our…
★ 869
NeumAI
Neum AI is a best-in-class framework to manage the creation and synchronization of vector embeddings at large…
★ 864
firecrawl-app-examples
🔥 This repository contains complete application examples, including websites and other projects, developed…
★ 782
swiftide
Fast, streaming indexing, query, and agentic LLM applications in Rust
★ 760
manuscript-core
Manuscript is a revolutionary blockchain data streaming framework. With Manuscript, you can seamlessly…
★ 690
bagofwords
Chat with your data - with memory, rules, and observability built in. Deploy in 2 minutes
★ 446
Data-Science-Hacks
Data Science Hacks consists of tips, tricks to help you become a better data scientist. Data science hacks…
★ 432
marimo-pair
Drop agents inside running marimo notebook sessions
★ 402
lionagi
An intelligence orchestra
★ 400
🔗 Famiglie affini

Misurato dai temi di GitHub condivisi da entrambi i progetti, ponderato in base a quanto è raro ciascun tema.