grpo

25 progetti condividono questo topic GitHub

grpo — ms-swift ★15kgrpoAI-Research-SKILLs — ★11.2kART — ★10.5kAgentGuide — ★7.6kVLM-R1 — ★6kreasoning-from-scratch — ★4.8khands-on-modern-rl — ★3.4kSkywork-R1V — ★3.2kverl-agent — ★2.2kMixGRPO — ★1.2kjudgeval — ★1kverl-omni — ★676VisualThinker-R1-Zero — ★624AutoVLA — ★610Open-AgentRL — ★595Awesome-RL-for-Video-Generation — ★574Relax — ★550ABC-GRPO — ★442Agentic-RAG-R1 — ★428LightRFT — ★404BioReason — ★399OpenThinkIMG — ★399agent_learning — ★342AlphaDrive — ★332DataClaw0 — ★117AI-Research-SKILLs★ 11.2kART★ 10.5kAgentGuide★ 7.6kVLM-R1★ 6kreasoning-from-scratch★ 4.8khands-on-modern-rl★ 3.4kSkywork-R1V★ 3.2kverl-agent★ 2.2kMixGRPO★ 1.2kjudgeval★ 1kverl-omni★ 676VisualThinker-R1-Zero★ 624AutoVLA★ 610Open-AgentRL★ 595Awesome-RL-for-Video-Gen…★ 574Relax★ 550ABC-GRPO★ 442Agentic-RAG-R1★ 428LightRFT★ 404BioReason★ 399OpenThinkIMG★ 399agent_learning★ 342AlphaDrive★ 332DataClaw0★ 117

Le linee collegano membri che sono misurabilmente correlati tra loro. La dimensione dei punti riflette le stelle.

🧬 Membri
ms-swift
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4,…
★ 15k
AI-Research-SKILLs
Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills…
★ 11.2k
ART
Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents…
★ 10.5k
AgentGuide
https://adongwanai.github.io/AgentGuide | AI Agent开发指南 | LangGraph实战 | 高级RAG |…
★ 7.6k
VLM-R1
Solve Visual Understanding with Reinforced VLMs
★ 6k
reasoning-from-scratch
Implement a reasoning LLM in PyTorch from scratch, step by step
★ 4.8k
hands-on-modern-rl
🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and…
★ 3.4k
Skywork-R1V
Skywork-R1V is an advanced multimodal AI model series developed by Skywork AI, specializing in…
★ 3.2k
verl-agent
verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the…
★ 2.2k
MixGRPO
[ECCV 2026] MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
★ 1.2k
judgeval
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and…
★ 1k
verl-omni
Multimodal RL training framework for diffusion & omni models
★ 676
VisualThinker-R1-Zero
Explore the Multimodal “Aha Moment” on 2B Model
★ 624
AutoVLA
[NeurIPS 2025] AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive…
★ 610
Open-AgentRL
RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic Scenarios
★ 595
Awesome-RL-for-Video-Generation
A curated list of papers on reinforcement learning for video generation
★ 574
Relax
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
★ 550
ABC-GRPO
Code For Adaptive-Boundary-Clipping GRPO. arxiv.org/pdf/2601.03895
★ 442
Agentic-RAG-R1
Agentic RAG R1 Framework via Reinforcement Learning
★ 428
LightRFT
LightRFT: Light, Efficient, Omni-modal & Reward-model Driven Reinforcement Fine-Tuning Framework
★ 404
BioReason
BioReason: Incentivizing Multimodal Biological Reasoning within a DNA-LLM Model | NeurIPS '25
★ 399
OpenThinkIMG
OpenThinkIMG is an end-to-end open-source framework that empowers LVLMs to think with images.
★ 399
agent_learning
A systematic AI Agent development tutorial covering LLM agents, RAG, tool use, memory systems, multi-agent…
★ 342
AlphaDrive
Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning
★ 332
DataClaw0
DataClaw: Agentic Tailoring Multimodal Data from Raw Streams — coming soon (code, weights, dataset &…
★ 117
🔗 Famiglie affini

Misurato dai temi di GitHub condivisi da entrambi i progetti, ponderato in base a quanto è raro ciascun tema.