rlhf

32 progetti condividono questo topic GitHub

rlhf — LlamaFactory ★74.5krlhfLLMSurvey — ★12.2kInternLM — ★7.3kChinese-LLaMA-Alpaca-2 — ★7.1kalignment-handbook — ★5.7kreasoning-from-scratch — ★5.1kargilla — ★5.1kalign-anything — ★4.7kawesome-RLHF — ★4.4khands-on-modern-rl — ★4.2kROLL — ★3.4kTorchLeet — ★2.5krlhf-book — ★2.2kalpaca_eval — ★2kAgentsMeetRL — ★1.8kImageReward — ★1.7ksafe-rlhf — ★1.6kWebGLM — ★1.6kRLHF-Reward-Modeling — ★1.5kSimPO — ★959verl-omni — ★915AlignLLMHumanSurvey — ★738Cornucopia-LLaMA-Fin-Chinese — ★656SPPO — ★589Relax — ★582Relax — ★580Awesome-LLM-On-Policy-Distillation — ★525dLLM-RL — ★520step_into_llm — ★480ABC-GRPO — ★442quick-start-guide-to-llms — ★395llm-flashcards — ★129LLMSurvey★ 12.2kInternLM★ 7.3kChinese-LLaMA-Alpaca-2★ 7.1kalignment-handbook★ 5.7kreasoning-from-scratch★ 5.1kargilla★ 5.1kalign-anything★ 4.7kawesome-RLHF★ 4.4khands-on-modern-rl★ 4.2kROLL★ 3.4kTorchLeet★ 2.5krlhf-book★ 2.2kalpaca_eval★ 2kAgentsMeetRL★ 1.8kImageReward★ 1.7ksafe-rlhf★ 1.6kWebGLM★ 1.6kRLHF-Reward-Modeling★ 1.5kSimPO★ 959verl-omni★ 915AlignLLMHumanSurvey★ 738Cornucopia-LLaMA-Fin-Chi…★ 656SPPO★ 589Relax★ 582 · GitHub ↗Relax★ 580Awesome-LLM-On-Policy-Di…★ 525dLLM-RL★ 520step_into_llm★ 480ABC-GRPO★ 442quick-start-guide-to-llm…★ 395llm-flashcards★ 129

Le linee collegano membri che sono misurabilmente correlati tra loro. La dimensione dei punti riflette le stelle.

🧬 Membri
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74.5k
LLMSurvey
The official GitHub page for the survey paper "A Survey of Large Language Models".
★ 12.2k
InternLM
Official release of InternLM series (InternLM, InternLM2, InternLM2.5, InternLM3).
★ 7.3k
Chinese-LLaMA-Alpaca-2
中文LLaMA-2 & Alpaca-2大模型二期项目 + 64K超长上下文模型 (Chinese LLaMA-2 & Alpaca-2 LLMs…
★ 7.1k
alignment-handbook
Robust recipes to align language models with human and AI preferences
★ 5.7k
reasoning-from-scratch
Implement a reasoning LLM in PyTorch from scratch, step by step
★ 5.1k
argilla
Argilla is a collaboration tool for AI engineers and domain experts to build high-quality datasets
★ 5.1k
align-anything
Align Anything: Training All-modality Model with Feedback
★ 4.7k
awesome-RLHF
A curated list of reinforcement learning with human feedback resources (continually updated)
★ 4.4k
hands-on-modern-rl
🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and…
★ 4.2k
ROLL
An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models
★ 3.4k
TorchLeet
LeetCode for PyTorch — 65 ML/AI interview problems from real interviews at Google, Meta, Anthropic. Jupyter…
★ 2.5k
rlhf-book
Textbook on reinforcement learning from human feedback
★ 2.2k
alpaca_eval
An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and…
★ 2k
AgentsMeetRL
Awesome List for Agentic RL
★ 1.8k
ImageReward
[NeurIPS 2023] ImageReward: Learning and Evaluating Human Preferences for Text-to-image Generation
★ 1.7k
safe-rlhf
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
★ 1.6k
WebGLM
WebGLM: An Efficient Web-enhanced Question Answering System (KDD 2023)
★ 1.6k
RLHF-Reward-Modeling
Recipes to train reward model for RLHF.
★ 1.5k
SimPO
[NeurIPS 2024] SimPO: Simple Preference Optimization with a Reference-Free Reward
★ 959
verl-omni
Multimodal RL training framework for diffusion & omni models
★ 915
AlignLLMHumanSurvey
Aligning Large Language Models with Human: A Survey
★ 738
Cornucopia-LLaMA-Fin-Chinese
聚宝盆(Cornucopia): 中文金融系列开源可商用大模型,并提供一套高效轻量化的垂直领…
★ 656
SPPO
The official implementation of Self-Play Preference Optimization (SPPO)
★ 589
Relax
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
★ 582 · GitHub ↗
Relax
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
★ 580
Awesome-LLM-On-Policy-Distillation
A curated collection of papers and resources on On-Policy Distillation for Large Language Models.
★ 525
dLLM-RL
[ICLR 2026] Official code for TraceRL: Revolutionizing post-training for Diffusion LLMs, powering the SOTA…
★ 520
step_into_llm
MindSpore online courses: Step into LLM
★ 480
ABC-GRPO
Code For Adaptive-Boundary-Clipping GRPO. arxiv.org/pdf/2601.03895
★ 442
quick-start-guide-to-llms
The Official Repo for "Quick Start Guide to Large Language Models"
★ 395
llm-flashcards
Visual knowledge bank for understanding large language models, with 180 concept cards from tokenization to…
★ 129
🔗 Famiglie affini

Misurato dai temi di GitHub condivisi da entrambi i progetti, ponderato in base a quanto è raro ciascun tema.