grpo

26 projets partagent ce topic GitHub

grpo — ms-swift ★15.5kgrpoART — ★10.5kAgentGuide — ★9kVLM-R1 — ★6kreasoning-from-scratch — ★5.1khands-on-modern-rl — ★4.2kSkywork-R1V — ★3.2kverl-agent — ★2.3kMixGRPO — ★1.2kjudgeval — ★1.1kverl-omni — ★915VisualThinker-R1-Zero — ★624AutoVLA — ★603Open-AgentRL — ★585Relax — ★582Relax — ★580Awesome-RL-for-Video-Generation — ★565agent_learning — ★475ABC-GRPO — ★442Agentic-RAG-R1 — ★426LightRFT — ★403BioReason — ★402OpenThinkIMG — ★397AlphaDrive — ★332NEWTON — ★143DataClaw0 — ★117ART★ 10.5kAgentGuide★ 9kVLM-R1★ 6kreasoning-from-scratch★ 5.1khands-on-modern-rl★ 4.2kSkywork-R1V★ 3.2kverl-agent★ 2.3kMixGRPO★ 1.2kjudgeval★ 1.1kverl-omni★ 915VisualThinker-R1-Zero★ 624AutoVLA★ 603Open-AgentRL★ 585Relax★ 582 · GitHub ↗Relax★ 580Awesome-RL-for-Video-Gen…★ 565agent_learning★ 475ABC-GRPO★ 442Agentic-RAG-R1★ 426LightRFT★ 403BioReason★ 402OpenThinkIMG★ 397AlphaDrive★ 332NEWTON★ 143DataClaw0★ 117

Les traits relient les membres réellement apparentés entre eux. La taille des points suit les étoiles.

🧬 Membres
ms-swift
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4,…
★ 15.5k
ART
Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents…
★ 10.5k
AgentGuide
https://adongwanai.github.io/AgentGuide | AI Agent开发指南 | LangGraph实战 | 高级RAG |…
★ 9k
VLM-R1
Solve Visual Understanding with Reinforced VLMs
★ 6k
reasoning-from-scratch
Implement a reasoning LLM in PyTorch from scratch, step by step
★ 5.1k
hands-on-modern-rl
🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and…
★ 4.2k
Skywork-R1V
Skywork-R1V is an advanced multimodal AI model series developed by Skywork AI, specializing in…
★ 3.2k
verl-agent
verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the…
★ 2.3k
MixGRPO
[ECCV 2026] MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
★ 1.2k
judgeval
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and…
★ 1.1k
verl-omni
Multimodal RL training framework for diffusion & omni models
★ 915
VisualThinker-R1-Zero
Explore the Multimodal “Aha Moment” on 2B Model
★ 624
AutoVLA
[NeurIPS 2025] AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive…
★ 603
Open-AgentRL
RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic Scenarios
★ 585
Relax
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
★ 582 · GitHub ↗
Relax
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
★ 580
Awesome-RL-for-Video-Generation
A curated list of papers on reinforcement learning for video generation
★ 565
agent_learning
A systematic AI Agent development tutorial covering LLM agents, RAG, tool use, memory systems, multi-agent…
★ 475
ABC-GRPO
Code For Adaptive-Boundary-Clipping GRPO. arxiv.org/pdf/2601.03895
★ 442
Agentic-RAG-R1
Agentic RAG R1 Framework via Reinforcement Learning
★ 426
LightRFT
LightRFT: Light, Efficient, Omni-modal & Reward-model Driven Reinforcement Fine-Tuning Framework
★ 403
BioReason
BioReason: Incentivizing Multimodal Biological Reasoning within a DNA-LLM Model | NeurIPS '25
★ 402
OpenThinkIMG
OpenThinkIMG is an end-to-end open-source framework that empowers LVLMs to think with images.
★ 397
AlphaDrive
Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning
★ 332
NEWTON
NEWTON: Agentic Planning for Physically Grounded Video Generation
★ 143
DataClaw0
DataClaw: Agentic Tailoring Multimodal Data from Raw Streams — coming soon (code, weights, dataset &…
★ 117
🔗 Familles voisines

Mesuré à partir des topics GitHub communs aux deux projets, pondérés par leur rareté.