rlhf

42 projetos partilham este topic do GitHub

rlhf — LlamaFactory ★73.6krlhfOpen-Assistant — ★37.4kLLMSurvey — ★12.2kInternLM — ★7.3kChinese-LLaMA-Alpaca-2 — ★7.1kalignment-handbook — ★5.7ktransformerlab-app — ★5.2kargilla — ★5.1kKiln — ★5kreasoning-from-scratch — ★4.8kalign-anything — ★4.7kawesome-RLHF — ★4.4khands-on-modern-rl — ★3.4kdistilabel — ★3.3kROLL — ★3.3krlhf-book — ★2.2kalpaca_eval — ★2kAgentsMeetRL — ★1.7kImageReward — ★1.7ksafe-rlhf — ★1.6kWebGLM — ★1.6kRLHF-Reward-Modeling — ★1.5kxtreme1 — ★1.2kSimPO — ★956AlignLLMHumanSurvey — ★742verl-omni — ★676Cornucopia-LLaMA-Fin-Chinese — ★659Open-AgentRL — ★595SPPO — ★590awesome-on-policy-distillation — ★572TextRL — ★564Relax — ★550dLLM-RL — ★511step_into_llm — ★480Awesome-LLM-On-Policy-Distillation — ★479LLM-RLHF-Tuning — ★453ABC-GRPO — ★442JarvisEvo — ★420pykoi — ★410quick-start-guide-to-llms — ★393VADER — ★315Open-Assistant★ 37.4kLLMSurvey★ 12.2kInternLM★ 7.3kChinese-LLaMA-Alpaca-2★ 7.1kalignment-handbook★ 5.7ktransformerlab-app★ 5.2kargilla★ 5.1kKiln★ 5kreasoning-from-scratch★ 4.8kalign-anything★ 4.7kawesome-RLHF★ 4.4khands-on-modern-rl★ 3.4kdistilabel★ 3.3kROLL★ 3.3krlhf-book★ 2.2kalpaca_eval★ 2kAgentsMeetRL★ 1.7kImageReward★ 1.7ksafe-rlhf★ 1.6kWebGLM★ 1.6kRLHF-Reward-Modeling★ 1.5kxtreme1★ 1.2kSimPO★ 956AlignLLMHumanSurvey★ 742verl-omni★ 676Cornucopia-LLaMA-Fin-Chi…★ 659Open-AgentRL★ 595SPPO★ 590awesome-on-policy-distil…★ 572TextRL★ 564Relax★ 550dLLM-RL★ 511step_into_llm★ 480Awesome-LLM-On-Policy-Di…★ 479LLM-RLHF-Tuning★ 453ABC-GRPO★ 442JarvisEvo★ 420pykoi★ 410quick-start-guide-to-llm…★ 393VADER★ 315

Linhas conectam membros que estão mensuravelmente relacionados entre si. O tamanho do ponto reflete estrelas.

🧬 Membros
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 73.6k
Open-Assistant
OpenAssistant is a chat-based assistant that understands tasks, can interact with third-party systems, and…
★ 37.4k
LLMSurvey
The official GitHub page for the survey paper "A Survey of Large Language Models".
★ 12.2k
InternLM
Official release of InternLM series (InternLM, InternLM2, InternLM2.5, InternLM3).
★ 7.3k
Chinese-LLaMA-Alpaca-2
中文LLaMA-2 & Alpaca-2大模型二期项目 + 64K超长上下文模型 (Chinese LLaMA-2 & Alpaca-2 LLMs…
★ 7.1k
alignment-handbook
Robust recipes to align language models with human and AI preferences
★ 5.7k
transformerlab-app
The open source research environment for AI researchers to seamlessly train, evaluate, and scale models from…
★ 5.2k
argilla
Argilla is a collaboration tool for AI engineers and domain experts to build high-quality datasets
★ 5.1k
Kiln
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data…
★ 5k
reasoning-from-scratch
Implement a reasoning LLM in PyTorch from scratch, step by step
★ 4.8k
align-anything
Align Anything: Training All-modality Model with Feedback
★ 4.7k
awesome-RLHF
A curated list of reinforcement learning with human feedback resources (continually updated)
★ 4.4k
hands-on-modern-rl
🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and…
★ 3.4k
distilabel
Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and…
★ 3.3k
ROLL
An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models
★ 3.3k
rlhf-book
Textbook on reinforcement learning from human feedback
★ 2.2k
alpaca_eval
An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and…
★ 2k
AgentsMeetRL
Awesome List for Agentic RL
★ 1.7k
ImageReward
[NeurIPS 2023] ImageReward: Learning and Evaluating Human Preferences for Text-to-image Generation
★ 1.7k
safe-rlhf
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
★ 1.6k
WebGLM
WebGLM: An Efficient Web-enhanced Question Answering System (KDD 2023)
★ 1.6k
RLHF-Reward-Modeling
Recipes to train reward model for RLHF.
★ 1.5k
xtreme1
Xtreme1 is an all-in-one data labeling and annotation platform for multimodal data training and supports 3D…
★ 1.2k
SimPO
[NeurIPS 2024] SimPO: Simple Preference Optimization with a Reference-Free Reward
★ 956
AlignLLMHumanSurvey
Aligning Large Language Models with Human: A Survey
★ 742
verl-omni
Multimodal RL training framework for diffusion & omni models
★ 676
Cornucopia-LLaMA-Fin-Chinese
聚宝盆(Cornucopia): 中文金融系列开源可商用大模型,并提供一套高效轻量化的垂直领…
★ 659
Open-AgentRL
RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic Scenarios
★ 595
SPPO
The official implementation of Self-Play Preference Optimization (SPPO)
★ 590
awesome-on-policy-distillation
A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of…
★ 572
TextRL
Implementation of ChatGPT RLHF (Reinforcement Learning with Human Feedback) on any generation model in…
★ 564
Relax
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
★ 550
dLLM-RL
[ICLR 2026] Official code for TraceRL: Revolutionizing post-training for Diffusion LLMs, powering the SOTA…
★ 511
step_into_llm
MindSpore online courses: Step into LLM
★ 480
Awesome-LLM-On-Policy-Distillation
A curated collection of papers and resources on On-Policy Distillation for Large Language Models.
★ 479
LLM-RLHF-Tuning
LLM Tuning with PEFT (SFT+RM+PPO+DPO with LoRA)
★ 453
ABC-GRPO
Code For Adaptive-Boundary-Clipping GRPO. arxiv.org/pdf/2601.03895
★ 442
JarvisEvo
[CVPR' 2026] JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator…
★ 420
pykoi
pykoi: Active learning in one unified interface
★ 410
quick-start-guide-to-llms
The Official Repo for "Quick Start Guide to Large Language Models"
★ 393
VADER
Video Diffusion Alignment via Reward Gradients. We improve a variety of video diffusion models such as…
★ 315
llm-flashcards
Visual knowledge bank for understanding large language models, with 180 concept cards from tokenization to…
★ 111
🔗 Familias relacionadas

Medido a partir dos tópicos do GitHub compartilhados por ambos os projetos, ponderado pela raridade de cada tópico.