rl

19 projets partagent ce topic GitHub

rl — Llama-Chinese ★14.7krldopamine — ★10.9kAReaL — ★5.7kEasyR1 — ★5.1khands-on-modern-rl — ★4.2kagent-sandbox — ★3.7kAwesome-RL-for-LRMs — ★2.5kall-rl-algorithms — ★1.9kPRIME — ★1.9kunderstand-r1-zero — ★1.3kSDPO — ★1.1kjudgeval — ★1.1ktensorlake — ★995zeroth-bot — ★799mobilegym — ★778meta-agents-research-environments — ★550DeepThinkVLA — ★527Agentic-RAG-R1 — ★426DiffusionOPSD — ★239dopamine★ 10.9kAReaL★ 5.7kEasyR1★ 5.1khands-on-modern-rl★ 4.2kagent-sandbox★ 3.7kAwesome-RL-for-LRMs★ 2.5kall-rl-algorithms★ 1.9kPRIME★ 1.9kunderstand-r1-zero★ 1.3kSDPO★ 1.1kjudgeval★ 1.1ktensorlake★ 995zeroth-bot★ 799mobilegym★ 778meta-agents-research-env…★ 550DeepThinkVLA★ 527Agentic-RAG-R1★ 426DiffusionOPSD★ 239

Les traits relient les membres réellement apparentés entre eux. La taille des points suit les étoiles.

🧬 Membres
Llama-Chinese
★ 14.7k
dopamine
Dopamine is a research framework for fast prototyping of reinforcement learning algorithms.
★ 10.9k
AReaL
The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
★ 5.7k
EasyR1
EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL
★ 5.1k
hands-on-modern-rl
🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and…
★ 4.2k
agent-sandbox
agent-sandbox enables easy management of isolated, stateful, singleton workloads, ideal for use cases like AI…
★ 3.7k
Awesome-RL-for-LRMs
A Survey of Reinforcement Learning for Large Reasoning Models
★ 2.5k
all-rl-algorithms
Implementation of all RL algorithms in a simpler way
★ 1.9k
PRIME
Scalable RL solution for advanced reasoning of language models
★ 1.9k
understand-r1-zero
Understanding R1-Zero-Like Training: A Critical Perspective
★ 1.3k
SDPO
Reinforcement Learning via Self-Distillation (SDPO)
★ 1.1k
judgeval
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and…
★ 1.1k
tensorlake
Tensorlake is a serverless runtime for sandboxes and deploying background agentic applications
★ 995
zeroth-bot
3D-printed open-source humanoid robot platform for sim-to-real and RL
★ 799
mobilegym
[EMNLP 2026] MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research ·…
★ 778
meta-agents-research-environments
Meta Agents Research Environments is a comprehensive platform designed to evaluate AI agents in dynamic,…
★ 550
DeepThinkVLA
DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
★ 527
Agentic-RAG-R1
Agentic RAG R1 Framework via Reinforcement Learning
★ 426
DiffusionOPSD
🔥 On-Policy Self-Distillation in Diffusion Models
★ 239
🔗 Familles voisines

Mesuré à partir des topics GitHub communs aux deux projets, pondérés par leur rareté.