A curated, opinionated catalogue of state-of-the-art tools, research, and companies for post-training and fine-tuning foundation models — covering instruction tuning, preference optimization, RLHF, RLVR, and agentic reinforcement learning. Organized around Nathan Lambert's RLHF Book (rlhfbook.com) as the canonical curriculum, then extended with industry tooling, US-university research, and the startup landscape.
Foundation-model post-training is the suite of techniques that turns a base model trained on next-token prediction into something useful: an instruction-follower, a chat assistant, a reasoner, a tool-using agent. The field has moved fast — from "SFT + simple RLHF" in 2023 to a dense ecosystem of preference optimization, verifiable rewards, process rewards, on-policy distillation, agentic RL, and synthetic-data flywheels in 2026. This list tries to be the map that wasn't there when the field exploded.
Scope. Open-source frameworks, reproducible recipes, benchmarks, datasets, key papers, and the companies (big tech, startups, non-profit labs) shaping the field. Heavy emphasis on US-based contributions but not exclusive.
Curation principle. Inclusion requires either a public artifact (code, paper, model, dataset, benchmark) or material commercial relevance. No vaporware.
Maintenance principle. Prefer durable technical descriptions over fast-decaying valuation, headcount, or revenue claims. When market details are included, add a verification date or move them to a separate landscape note.
- How to Use This List
- Foundational Reading
- The Canonical Post-Training Recipe
- Stage 1 — Instruction Fine-Tuning (SFT)
- Stage 2 — Preference Data & Reward Modeling
- Stage 3 — Direct Alignment Algorithms (DPO and family)
- Stage 4 — Reinforcement Learning (PPO, GRPO, REINFORCE)
- Stage 5 — Reinforcement Learning with Verifiable Rewards (RLVR)
- Stage 6 — Agentic / Multi-Turn / Tool-Use RL
- Stage 7 — On-Policy Distillation and Model Merging
- Stage 8 — Character, Personality, and Product Training
- Cross-Cutting — Synthetic Data Generation
- Cross-Cutting — Evaluation
- Cross-Cutting — Interpretability and Behavior Analysis
- Open Models with Fully Open Post-Training Recipes
- Training Frameworks (SFT-focused)
- Training Frameworks (RL-focused)
- Serving and Multi-Adapter Infrastructure
- Research Labs and University Groups
- Startups and Companies
- Big Tech Foundation Labs
- Courses, Blogs, and Newsletters
- Conference Tutorials
- Open Questions
This list is intentionally structured as both a learning path and an implementation map.
- Learning path: Foundational Reading → SFT → Preference Data → DPO → RLHF / RLVR → Agentic RL.
- Practitioner path: Datasets → Training Frameworks → Evaluation → Serving and Multi-Adapter Infrastructure.
- Founder / enterprise path: Managed platforms → Data operations → Evaluation, safety, and governance → Serving.
- Research path: Open recipes → University groups → Open questions.
Suggested tags used throughout:
[paper]— research paper or technical report.[code]— runnable implementation, framework, or library.[dataset]— training, preference, reward, or evaluation dataset.[benchmark]— evaluation benchmark or leaderboard.[platform]— managed or commercial system relevant to fine-tuning / post-training.[recipe]— reproducible training pipeline, model card, or training mixture.
The texts to read before anything else.
- The RLHF Book — Nathan Lambert (Ai2). The authoritative, continuously-updated reference. Covers history, instruction tuning, reward modeling, rejection sampling, policy gradients, DPO, RLVR, reasoning, tool use, character training, evaluation, and overoptimization. GitHub | arXiv | Manning | Code library | Course. Includes reference implementations of PPO, REINFORCE, GRPO, RLOO, preference RMs, ORMs, PRMs, and DPO variants.
- Spinning Up in Deep RL — Josh Achiam (OpenAI). The canonical RL primer the RLHF Book points to for prerequisites.
- Reinforcement Learning: An Introduction — Sutton & Barto, 2nd ed. The classical reference.
- Illustrating RLHF — Lambert, Castricato, von Werra, Havrilla (Hugging Face, 2022). The blog post that introduced the modern formulation to a wide audience.
- Training language models to follow instructions with human feedback — Ouyang et al. (OpenAI, 2022). The InstructGPT paper. The ChatGPT-era starting point.
- Constitutional AI — Bai et al. (Anthropic, 2022). RLAIF and the first scalable principle-based alignment recipe.
- Training a Helpful and Harmless Assistant with RLHF — Bai et al. (Anthropic, 2022). The HH preference data and the canonical "helpful + harmless" framing.
- Tülu 3: Pushing Frontiers in Open Language Model Post-Training — Ai2 (Lambert et al., 2024). The most complete public post-training recipe. Introduced RLVR at scale.
The Lambert formulation, with chapter cross-references to The RLHF Book:
- Instruction Fine-Tuning (SFT) — Ch. 4. Teach the chat format and basic instruction following.
- Preference Data Collection — Ch. 6. Build preference pairs (or rubrics, or verifiable signals).
- Reward Modeling — Ch. 5/7. Train a scalar reward model on preference data; optionally outcome (ORM) or process (PRM).
- Rejection Sampling — Ch. 9/10. Generate many completions, keep the best by RM, SFT on them. The cheapest RLHF.
- Direct Alignment — Ch. 8/12. DPO and family — optimize the RL objective directly from preference pairs.
- Policy-Gradient RL — Ch. 6/11. PPO, GRPO, REINFORCE++. The heaviest but generally most effective stage.
- RLVR — Verifiable rewards for math, code, structured tasks. The 2024-2026 renaissance.
- Character/Product Training — Ch. 17. Synthetic-data-heavy persona shaping. The "tricks" everyone uses, nobody publishes.
- Evaluation — Ch. 18/19. Continuous, multi-axis, with regression gates against general capability benchmarks.
Modern recipes mix-and-match these stages; the Tülu 3 sequence (SFT → DPO → RLVR) is currently the strongest public baseline.
- Finetuned Language Models Are Zero-Shot Learners (FLAN) — Wei et al. (Google, 2021). The progenitor of instruction tuning at scale.
- Self-Instruct: Aligning Language Models with Self-Generated Instructions — Wang et al. (UW + Ai2, 2022). The synthetic-instruction-data starting gun.
- LIMA: Less Is More for Alignment — Zhou et al. (Meta, 2023). 1,000 well-curated examples beat 50,000 mediocre ones.
- Scaling Instruction-Finetuned Language Models (FLAN-T5) — Chung et al. (Google, 2022).
- Alpaca — Taori et al. (Stanford CRFM, 2023). The Self-Instruct + Llama recipe that catalyzed the open community.
- Tülu 3 SFT Mixture (Ai2) — current open-recipe gold standard.
- OpenHermes 2.5 (Nous Research) — 1M curated SFT examples.
- No Robots (Hugging Face) — 10K human-written, no AI assistance.
- Dolly 15K (Databricks) — fully human-generated.
- OASST1 / OASST2 (LAION) — community-generated conversation trees.
- Deep RL from Human Preferences — Christiano et al. (OpenAI + DeepMind, 2017). The foundational paper. Pre-LLM, on Atari.
- Anthropic HH-RLHF — Bai et al. (2022). Helpfulness/harmlessness preference data, still a benchmark.
- Scaling Laws for Reward Model Overoptimization — Gao, Schulman, Hilton (OpenAI, 2022). Why reward hacking is the problem you'll spend most of your time on.
- RewardBench — Lambert et al. (Ai2, 2024). The first systematic reward-model benchmark.
- Self-Taught Evaluator — Wang et al. (Meta, 2024). Reward models trained on synthetic preference data.
- Process Reward Models (PRMs) — Let's Verify Step by Step — Lightman et al. (OpenAI, 2023). Step-level supervision for reasoning.
- Generative Reward Models — Mahan et al. (SynthLabs + Stanford, 2024). LLM-as-judge formalized.
- UltraFeedback — Cui et al. The most-used open preference dataset.
- Skywork-Reward-Preference-80K — high-quality curated preferences.
- HelpSteer2 (NVIDIA) — multi-attribute preference data.
- PRM800K (OpenAI) — process-level math reasoning labels.
- Nectar (Berkeley Starling) — 7-way preferences from GPT-4-as-judge.
- RewardBench (Ai2) — RM evaluation harness.
- Argilla
[code] [data]— human and AI feedback workflows for preference data, evaluation sets, and dataset review. - Label Studio
[code] [data]— general-purpose labeling and review platform that can support preference, safety, and task-specific annotation. - Lilac
[code] [data]— dataset exploration, clustering, search, and curation for LLM training and evaluation data. - Cleanlab
[code] [data-quality]— data-quality tooling useful for finding label issues, outliers, and low-confidence examples before fine-tuning.
- Direct Preference Optimization — Rafailov et al. (Stanford, 2023). The paper that broke the field open in late 2023 / early 2024. Lambert credits Zephyr-Beta, Tülu 2 and the lower-learning-rate "trick" for making DPO actually work in practice.
- Identity Preference Optimization (IPO) — Azar et al. (DeepMind, 2023).
- KTO: Model Alignment as Prospect Theoretic Optimization — Ethayarajh et al. (Stanford + Contextual, 2024). Works on unpaired binary feedback.
- ORPO: Monolithic Preference Optimization without Reference Model — Hong et al. (KAIST, 2024). SFT and preference learning in one stage.
- SimPO: Simple Preference Optimization with a Reference-Free Reward — Meng et al. (Princeton + UVa, 2024).
- Iterative DPO / sDPO — multiple groups. On-policy data refresh during DPO.
- The Alignment Handbook (Hugging Face) — practical DPO/SFT recipes. The reproducible Zephyr recipe lives here.
- Proximal Policy Optimization (PPO) — Schulman et al. (OpenAI, 2017). The workhorse.
- Group Relative Policy Optimization (GRPO) — DeepSeek (DeepSeekMath, 2024). No critic model, group-relative advantages. The DeepSeek-R1 enabler.
- REINFORCE++ / RLOO — Ahmadian et al. (Cohere, 2024). REINFORCE leave-one-out — simpler and often competitive with PPO.
- Secrets of RLHF in Large Language Models (PPO tricks) — Zheng et al. (Fudan, 2023).
- The N+ Implementation Details of RLHF (Hugging Face / Costa Huang). Critical practitioner reference.
- DAPO: Decoupled Clip and Dynamic Sampling — ByteDance Seed (2025). GRPO improvements widely adopted.
- ScaleRL (Meta, 2025) — validation of REINFORCE++-baseline at large scale.
The book ships canonical implementations of PPO, REINFORCE, GRPO, and RLOO in code/policy_gradients/.
The 2024-2026 renaissance. Replace the reward model with a programmatic verifier where ground truth exists.
- Tülu 3 / RLVR — Ai2 (Lambert et al., 2024). The paper that introduced RLVR as a named technique and scaled it. Gains of 1.7 / 3.3 / 1.3 points over DPO checkpoints on MATH / GSM8K / IFEval; scaled better at 405B than smaller scales.
- DeepSeek-R1 — DeepSeek (2025). GRPO + RLVR for reasoning. The proof point.
- OpenAI o1 system card — OpenAI (2024). The closed-source precursor.
- Let's Verify Step by Step — Lightman et al. (OpenAI, 2023). PRMs for math.
- Reinforcement Pre-Training (RPT) — multiple groups. RL applied during/before SFT, not after.
- Golden Goose (2026) — synthesize unlimited RLVR tasks from raw internet text; 4B model trained on cyber-domain data surpasses 7B domain-specialist.
- open-instruct (Ai2) — the canonical RLVR codebase. Used to train Tülu 3 and OLMo 2/3.
- verl — production-grade GRPO/PPO training, used heavily for reasoning RL.
- VeRL-Pipeline (HybridFlow paper).
The current frontier. Multi-step, tool-using, long-horizon agents trained with RL on trajectories, not single completions.
- AgentFlow / Flow-GRPO — Stanford (ICLR 2026 Oral). In-the-flow planner training: 17.2% gain with online RL vs. 19.0% collapse with SFT — strongest published evidence that SFT actively hurts on agentic tasks and RL is required.
- Self-Challenging Agents — Meta + UC Berkeley (2025). RL on self-generated synthetic tasks: 95.8% relative improvement on Llama-3.1-8B across four tool-use environments.
- SWiRL: Step-Wise Reinforcement Learning — Stanford + Google DeepMind (Goldie, Mirhoseini, Manning, 2025). 21.5%, 12.3%, 14.8%, 11.1%, 15.3% gains on GSM8K, HotPotQA, CofCA, MuSiQue, BeerQA; cross-task generalization (training on HotPotQA improves GSM8K by 16.9%).
- Verlog — CMU (2025). Multi-turn RL framework for episodes >400 turns. The state-of-the-art for long-horizon training.
- MUA-RL — Multi-turn user-agent RL with LLM-simulated users in the training loop.
- CodeAct — UIUC + Apple (2024). Executable code as the action format with execution feedback. Foundational for code agents.
- SWE-agent — Princeton (2024). Agent-computer interface for software engineering.
- τ-bench — Shunyu Yao et al. (Princeton, 2024). Tool-agent-user benchmark with the pass^k reliability metric.
- VerlTool (2025). Holistic agentic RL with comprehensive tool support — useful comparison table of frameworks.
- Agentic Context Engineering (ACE) — Stanford + UC Berkeley + SambaNova (2025). Self-improvement without weight updates; >10pp gain on AppWorld, 8.6% on financial reasoning.
- OpenHands (CMU / Neubig group). The open-source agent framework with most mindshare.
- Verlog (CMU) — long-horizon training framework.
- Torchforge (Meta PyTorch) — production-style agentic RL.
- OpenEnv (Meta PyTorch) — standard environment interface for agent RL.
- VerlTool — verl extended with comprehensive tool support.
- NeMo-Gym (NVIDIA) — agent-based RLHF with external evaluation environments; integrates with OpenRLHF.
- SWE-bench / Verified / Multimodal / Lite (Princeton) — the de facto coding-agent benchmark. 500 verified instances, Docker-based harness, Modal-based parallel evaluation. Used by every major lab.
- τ-bench (Sierra Research) — dynamic user-agent interaction.
- TheAgentCompany (CMU) — multi-role enterprise agent benchmark; 3,000 hours of researcher labor to build.
- WebArena (CMU) — realistic web agent benchmark.
- AgentBench — multi-environment agent eval.
- InterCode (Princeton) — interactive coding tasks.
- On-Policy Distillation of Language Models — Agarwal et al. (Google, 2023). Foundational.
- Distilling Step-by-Step — Hsieh et al. (Google + UW, 2023). Distill rationales, not just answers.
- MiniLLM — Microsoft (2023). Reverse-KL distillation for LLMs.
- Zephyr 7B — Hugging Face (2023). One of the first showcases of distillation + DPO on a small model.
- MergeKit (Arcee AI) — production merging library.
- TIES-Merging — Yadav et al. (UNC, 2023).
- DARE — Yu et al. (Microsoft + Princeton, 2023).
- SLERP — geometric merge; widely used baseline.
- Evolutionary Model Merge — Sakana AI (2024). Automated search over merge configurations.
The Lambert "tricks chapter" — what frontier labs actually do to round out models. The RLHF Book Ch. 17 expanded significantly in the print edition.
- Claude's Character (Anthropic, 2024). The canonical public statement on character training.
- Persona Vectors and persona subnetworks — additive (activation-space) and multiplicative (weight-space) persona interventions, no gradient updates.
- Sycophancy in Language Models — Sharma et al. (Anthropic, 2023). The cautionary tale.
- The RLHF Book Ch. 17 — the most complete public treatment.
Arguably the most under-credited multiplier in the modern stack.
- Self-Instruct — UW + Ai2 (2022). Bootstrap instructions from a teacher model.
- Evol-Instruct (WizardLM) — Microsoft + Peking (2023). Difficulty evolution of seed prompts.
- OSS-Instruct (Magicoder) — UIUC + Tsinghua (2023). Generate code instructions from seed source files.
- Persona Hub — Tencent (2024). 1B personas for diverse data generation.
- Nemotron-4 340B Technical Report — NVIDIA (2024). Heavily synthetic alignment data.
- SynthRL — verifiable visual reasoning data synthesis (ICML 2025).
- Golden Goose — synthesize RLVR tasks from unverifiable internet text.
- distilabel (Argilla / Hugging Face) — production synthetic-data pipelines.
- Magpie — self-aligning data generation from instruct models.
- Curator (Bespoke Labs) — production synthetic dataset creation.
- DataDreamer — synthetic-data generation, instruction-data workflows, and reproducible LLM data pipelines.
- DSPy (Stanford NLP) — programmatic prompt / data / pipeline optimization; useful for generating and optimizing task-specific training traces.
- MMLU / MMLU-Pro — broad knowledge.
- GSM8K / MATH — math reasoning.
- HumanEval / BigCodeBench / LiveCodeBench — code.
- IFEval — instruction following.
- BFCL (Berkeley Function Calling Leaderboard) — tool/function calling.
- GPQA — graduate-level science.
- ARC-AGI — abstraction and reasoning.
- Chatbot Arena (LMSYS / now Arena Intelligence) — human preference Elo across 6M+ votes.
- AlpacaEval 2.0 (Stanford). LLM-judge, length-controlled.
- MT-Bench (LMSYS) — multi-turn LLM-judge.
- Arena-Hard (LMSYS) — automated, hard prompts.
- WildBench (Ai2) — real user prompts.
- SWE-bench Verified (Princeton + OpenAI Preparedness).
- τ-bench (Sierra Research).
- TheAgentCompany (CMU).
- AgentBench, WebArena, OSWorld, BrowseComp (OpenAI).
- AppWorld
[benchmark]— benchmark for agents operating across realistic app APIs and user goals. - BrowserGym
[benchmark] [code]— browser-agent environments and evaluation tasks. - MiniWoB++
[benchmark]— classic web-interaction tasks for UI agents.
- Inspect AI (UK AI Safety Institute / now AI Security Institute). Probably the most adopted modern eval framework. Schema-compatible with Docent.
- lm-evaluation-harness (EleutherAI). Long-standing reference for capability eval.
- LightEval (Hugging Face).
- OpenAI Evals.
- HELM (Stanford CRFM). Holistic evaluation.
- Vals AI — third-party eval platform with public leaderboards.
- Epoch AI Benchmarks — independently reproduced frontier evals.
- NIST AI Risk Management Framework — Generative AI Profile
[governance]— risk-management reference for generative AI systems. - MLCommons AILuminate
[benchmark] [safety]— safety benchmark family for assessing model behavior across hazard categories. - OWASP Top 10 for LLM Applications
[security]— practical taxonomy for prompt injection, data leakage, model theft, supply-chain risks, and other LLM-app threats. - HarmBench
[benchmark] [safety]— standardized harmful-behavior evaluation for LLMs. - WMDP
[benchmark] [safety]— hazardous-knowledge evaluation for biosecurity, cybersecurity, and chemical domains. - CyberSecEval
[benchmark] [security]— cybersecurity-oriented model safety and capability evaluations.
- General capability should not collapse after SFT / DPO / RL.
- Instruction following should improve without overfitting to template artifacts.
- Safety refusals should improve without excessive false refusals.
- Reward models and LLM judges should be audited for verbosity, sycophancy, and style bias.
- Agentic training should report trajectory-level success, pass^k reliability, tool-call validity, and failure modes.
- Docent (Transluce, ex-Berkeley/MIT). Agent transcript analysis — clustering, rubrics, counterfactual intervention, "reward hacks across training steps" monitoring. The standard for agentic behavior analysis. GitHub.
- Goodfire Ember — interpretability platform for model-behavior analysis, steering, training-phase introspection, and production monitoring.
- Neuronpedia — community SAE/feature visualization.
- SAELens — sparse-autoencoder library.
- TransformerLens — mechanistic interpretability workbench.
- Towards Monosemanticity — Anthropic (2023). SAE-based interpretability foundation.
- Scaling Monosemanticity — Anthropic (2024). SAEs on Claude 3 Sonnet.
- Gemma Scope — Google DeepMind (2024). Open-sourced SAE suite.
The shortlist of models that release weights + data + training code + recipes.
- OLMo 2 / OLMo 3 (Ai2). 7B / 13B / 32B. Fully open, including Tülu-3 post-training pipeline.
- Tülu 3 / Tülu 3 405B (Ai2). The reference open recipe. Tülu 3 SFT mixture, DPO data, RLVR setup all released.
- Llama Instruct family (Meta). Weights open, recipe partially documented. Lambert was directly involved in some of these.
- Qwen Instruct family (Alibaba). Strong open recipes for chat and reasoning.
- DeepSeek-R1 / DeepSeek-V3 (DeepSeek). The RLVR-trained reasoning models. Recipe partially open.
- Zephyr 7B (Hugging Face). The Alignment Handbook reproducible reference.
- SmolLM3 (Hugging Face). Small-model post-training reference.
- Nemotron families (NVIDIA). Heavily synthetic-data post-training.
- Phi-4 (Microsoft). Synthetic-data-led small model recipe.
| Framework | Strength | Best for |
|---|---|---|
| Axolotl | YAML configs, broad model support, multi-GPU via FSDP2 | Production multi-GPU SFT/RLHF |
| Unsloth | Custom CUDA kernels, single-GPU speed, GRPO at 5GB VRAM | Single-GPU experimentation |
| Torchtune (PyTorch) | PyTorch-native, lean, QAT support | Pure-PyTorch shops |
| LLaMA-Factory | Broad UI/model support | Quick experimentation |
| Hugging Face TRL | Post-training library for SFT, DPO, PPO, GRPO, reward modeling, and related methods | Anyone in HF ecosystem |
| Hugging Face PEFT | LoRA, QLoRA, IA3, prefix tuning, and adapter-style parameter-efficient fine-tuning | Adapter-first workflows |
| bitsandbytes / QLoRA | 8-bit / 4-bit quantization and low-memory fine-tuning | Single-GPU and cost-sensitive fine-tuning |
| Accelerate | Lightweight distributed training launcher and device abstraction | Multi-GPU without heavy infra |
| DeepSpeed | ZeRO optimization, distributed training, memory efficiency | Full fine-tuning at scale |
| Liger Kernel | Triton kernels for memory-efficient LLM training | Throughput and memory optimization |
| Levanter (Stanford CRFM) | JAX/TPU, reproducible, named-axis tensors | TPU users |
| NVIDIA NeMo | Production scale, Megatron-LM under the hood | NVIDIA-stack enterprises |
| ColossalAI | Distributed training infra | Multi-node scaling |
| xtuner (Shanghai AI Lab) | Wide model coverage | Asia-developer-heavy |
The 2024-2026 RL-framework explosion. From the VerlTool comparison:
| Framework | Origin | Style | Notable |
|---|---|---|---|
| verl | ByteDance Seed | Sync, Ray + FSDP/Megatron | HybridFlow paper; widely adopted production-grade |
| OpenRLHF | Community-led | Sync/async, Ray + vLLM + ZeRO-3 | Scalable PPO / DPO / GRPO / REINFORCE-style RLHF and agentic RL workflows |
| TRL | Hugging Face | Sync, single-node to small distributed | Most-used learning entry point for SFT, DPO, PPO, GRPO, and reward modeling |
| open-instruct | Ai2 | Sync, Tülu/OLMo backbone | The Tülu 3 / RLVR reference |
| AReaL | Ant Research | Fully async | Scales to long-horizon reasoning |
| ROLL | Alibaba | Async | Puzzle-environment focus |
| slime | THUDM | Async | Emerging |
| Rlinf | SAIL Singapore | Async | |
| RL2 | Open | Async | |
| NeMo-Aligner + NeMo-Gym | NVIDIA | Sync | Production with environment integration |
| Torchforge + OpenEnv | Meta PyTorch | Async | Agentic-first, env standard |
| VerlTool | TIGER-Lab | Sync | verl extended for tool-use |
| RAGEN | RAGEN-AI / CMU | Multi-turn | Long-horizon agent training |
| trlX (legacy) | CarperAI | Sync | Pre-2024 reference; NeMo-backed for large scale |
Critical for any commercial post-training stack — serve many fine-tuned variants from one base.
- vLLM (Berkeley Sky Lab → vLLM project). PagedAttention; the de facto open inference engine. Native LoRA hot-swap.
- SGLang — high-throughput inference with strong RL integration.
- TensorRT-LLM (NVIDIA).
- Text Generation Inference (TGI) (Hugging Face).
- LoRAX (Predibase) — multi-LoRA serving from a single base.
- Punica — efficient multi-tenant LoRA.
- S-LoRA — Berkeley. Foundational paper on scalable LoRA serving.
- Percy Liang / CRFM — HELM, Alpaca, Levanter; benchmark and methodology anchor.
- Christopher Manning's group (Stanford NLP) — DSPy origin, SWiRL co-author.
- Stefano Ermon — DPO co-author.
- Tatsunori Hashimoto — AlpacaEval, KTO co-author.
- Azalia Mirhoseini — SWiRL, agentic systems.
- CS336 — Language Modeling from Scratch (course). The de facto modern curriculum.
- Sky Computing Lab (Ion Stoica, Joseph Gonzalez) — vLLM, SkyPilot, Ray, Chatbot Arena. Most prolific research-to-startup pipeline in the field.
- Jacob Steinhardt — Transluce co-origin; model behavior and evaluation.
- Sergey Levine — offline RL foundations applicable to LLMs.
- Berkeley Starling Team — Nectar, Starling reward models.
- Graham Neubig — OpenHands, TheAgentCompany; agent systems and evaluation.
- Aditi Raghunathan — robustness, post-training generalization.
- CMU Advanced NLP (Spring 2025) — uses OpenRLHF as teaching framework.
- Verlog / RAGEN-AI — long-horizon agent training.
- Princeton Language and Intelligence (PLI).
- Karthik Narasimhan / Shunyu Yao — SWE-bench, SWE-agent, τ-bench; the agent benchmark lineage.
- Tri Dao — FlashAttention; the inference enabler everyone uses.
- Sanjeev Arora — theoretical foundations.
- Jacob Andreas — MAIA, FIND; interpretability agents; Transluce co-origin.
- Antonio Torralba — vision-language and multi-modal.
- Dylan Hadfield-Menell (Algorithmic Alignment Group) — alignment-relevant fine-tuning.
- The non-profit benchmark for open post-training. Nathan Lambert, Hanna Hajishirzi, Noah Smith, Yejin Choi, Pradeep Dasigi.
- Tülu 3, OLMo 2/3, RewardBench, open-instruct, WildBench.
- Cornell — Alexander Rush, Claire Cardie; reasoning, agent training.
- Harvard — Stuart Shieber; test-time interaction scaling collaborations with CMU.
- Columbia — SWE-Bench-CL continual learning evaluation.
- UIUC — CodeAct (Heng Ji's group).
- UCLA — Kai-Wei Chang; agent reasoning.
- University of Washington — Yejin Choi (now at Stanford), Hanna Hajishirzi (Ai2), Tengyang Xie (Markov-state framing of RL post-training).
- NYU — Kyunghyun Cho, Sam Bowman (now Anthropic); preference learning theory.
- Reflection AI — Berkeley / DeepMind-rooted company centered on RL post-training for autonomous agents. A thesis-level competitor to watch.
- Imbue (formerly Generally Intelligent). Berkeley-rooted work on robust reasoning, coding agents, and agentic systems.
- Adept (now partially absorbed into Amazon AGI). Agent foundation models.
- Predibase — managed fine-tuning, RFT / GRPO workflows, and LoRAX multi-LoRA serving.
- Together AI — broad fine-tuning and inference platform with support for customization workflows across open models.
- Fireworks AI — inference and fine-tuning platform with an agent / reinforcement-fine-tuning orientation.
- Anyscale — Ray-based ML infrastructure and managed compute for large-scale training and serving.
- Modal — serverless GPU infrastructure; widely used for evaluation, batch inference, and ML workloads.
- Databricks + MosaicML — enterprise data + model-training stack for fine-tuning and domain adaptation.
- Snowflake — enterprise data platform with model customization and AI application workflows.
- OpenAI Fine-Tuning
[platform]— supervised fine-tuning and reinforcement fine-tuning workflows for OpenAI models. - Amazon Bedrock Model Customization
[platform]— managed fine-tuning, continued pre-training, distillation, model import, and customization workflows across supported foundation models. - Google Vertex AI Model Tuning
[platform]— managed tuning workflows for Gemini and model-garden models. - Azure AI Foundry / Azure OpenAI Fine-Tuning
[platform]— Azure-managed fine-tuning and deployment workflow for supported OpenAI models. - Databricks Mosaic AI Model Training
[platform]— enterprise fine-tuning and continued-training workflows on lakehouse data. - Hugging Face AutoTrain
[platform] [code]— low-code training workflow for fine-tuning LLMs and other model types.
- Transluce — non-profit. Docent (transcript analysis), Monitor (interpretability). Used to evaluate Claude 4, by governments for risk assessment.
- Goodfire — interpretability platform for model behavior analysis, feature steering, and production monitoring.
- Patronus AI — eval/safety platform.
- Arize — observability with LLM extensions.
- Cognition AI (Devin). Acquired Windsurf in 2025. Hundreds of thousands of merged PRs.
- Cursor (Anysphere) — AI-native coding environment and downstream consumer of code-agent advances.
- Sierra — customer-support agents; created τ-bench as a realistic user-agent benchmark.
- Harvey — legal AI agents and workflows.
- Decagon — customer support.
- Glean — enterprise search/agents.
- Hippocratic AI — clinical agents.
- Crescendo / Ada — CX agents.
- LangChain — agent / LLM application framework company behind LangChain, LangGraph, and LangSmith.
- LlamaIndex — RAG / agent framework.
- CrewAI — multi-agent orchestration.
- Pydantic AI (Pydantic / Samuel Colvin) — typed agent framework.
- Scale AI — preference data at scale (Meta acquired majority stake in 2025).
- Surge AI — high-quality human feedback.
- Argilla (now Hugging Face) — data annotation; distilabel maintainer.
- Bespoke Labs — Curator framework.
- Snorkel AI — programmatic data + alignment.
- Baseten — ML deployment.
- Replicate — model hosting.
- Hugging Face — TRL, Alignment Handbook, datasets, Inference Endpoints, distilabel via Argilla. The hub.
- Anyscale, Databricks (older AMPLab spinout), Arena Intelligence (Chatbot Arena spinout) — examples of Berkeley Sky Lab / AMPLab-style research-to-startup transfer.
- Anthropic — Claude family. Constitutional AI, character training, RLAIF lineage. Interpretability publishes via transformer-circuits.pub.
- OpenAI — GPT family. RLHF popularizer; o-series RLVR.
- Google DeepMind — Gemini family. SWiRL, Gemma open models, Gemma Scope.
- Meta AI / FAIR — Llama family. Llama Instruct recipes, ScaleRL, Self-Challenging Agents, Torchforge.
- Microsoft AI — Phi series; synthetic-data leadership; large investment in OpenAI.
- NVIDIA AI — Nemotron, NeMo, NeMo-Aligner, NeMo-Gym, HelpSteer datasets, TensorRT-LLM.
- Apple — on-device models; CodeAct co-authorship.
- Amazon AGI / Bedrock — Nova family; Bedrock AgentCore.
- xAI — Grok family.
- Stanford CS336 — Language Modeling from Scratch.
- The RLHF Book Course — Lambert's own lecture series.
- CMU Advanced NLP (Neubig).
- Berkeley CS 294/194-280 LLM Agents MOOC.
- Hugging Face NLP Course — practical orientation.
- Interconnects — Nathan Lambert. The single best running commentary on post-training.
- Hugging Face Blog — Costa Huang, Lewis Tunstall, Philipp Schmid; deep practical posts on TRL, alignment, evaluation.
- Anthropic Research — Claude character, persona vectors, interpretability.
- OpenAI Blog.
- Lilian Weng's Blog (lilianweng.github.io) — though less frequent post-2024.
- Cameron R. Wolfe — Deep (Learning) Focus — survey-quality writeups.
- Sebastian Raschka — Ahead of AI — practical and educational.
- The Gradient.
- Sky Computing Lab Blog.
- The Cognitive Revolution — Nathan Labenz; technical depth.
- Latent Space — Swyx / Alessio.
- Dwarkesh Podcast — deep technical interviews.
- No Priors — Sarah Guo / Elad Gil.
Curated tutorials from top AI/NLP/ML venues covering RLHF, alignment, post-training, and related topics. Sorted by recency within each venue.
- Reinforcement Learning from Human Feedback: Progress and Challenges — Ziegler, Leike, Stiennon (OpenAI), NeurIPS 2023. The definitive tutorial from the team that built the original RLHF pipeline. Covers the full stack from reward modeling to PPO.
- Efficient Training and Inference for Large Language Models — Beidi Chen, Tri Dao, NeurIPS 2023. FlashAttention, quantization, and efficient fine-tuning infrastructure.
- Advances in Reward Modeling — Lambert, Zhu (Ai2 / Stanford), NeurIPS 2024. RewardBench, generative RMs, and overoptimization.
- Deep Reinforcement Learning from Human Feedback — NeurIPS 2024. End-to-end RLHF with PPO and direct alignment alternatives.
- Instruction Tuning for Large Language Models — Jason Wei, Hyung Won Chung (Google Brain), ICML 2023. FLAN lineage and scaling instruction-following.
- Reinforcement Learning for Language Models — John Schulman (OpenAI), ICML 2023. PPO mechanics and the RLHF objective derivation from first principles.
- Aligning Language Models to Follow Instructions — OpenAI team, ICLR 2023. InstructGPT retrospective and lessons from RLHF at scale.
- Post-Training for Foundation Models — Nathan Lambert (Ai2), ICLR 2024. Comprehensive walkthrough of the modern SFT → DPO → RLVR pipeline; companion to the RLHF Book.
- Data-Efficient Fine-Tuning — Chelsea Finn, Percy Liang (Stanford), ICLR 2024. LIMA, data selection, and quality-over-quantity principles.
- Preference Optimization: Theory and Practice — Ryan Rafailov, Archit Sharma (Stanford), ICLR 2024. DPO derivation, failure modes, and best-practice hyperparameters.
- RLVR and Verifiable Reward Signals — DeepSeek + Ai2 authors, ICLR 2025. GRPO, process rewards for math and code, and the DeepSeek-R1 recipe.
- Agentic AI: Training, Evaluation, and Deployment — Karthik Narasimhan, Shunyu Yao (Princeton), ICLR 2025. SWE-bench, τ-bench, and multi-turn RL for tool-using agents.
- Instruction-Following and Human Feedback in NLP — Wang et al. (UW + Ai2), ACL 2023. Self-Instruct, preference data collection, and evaluation for instruction-tuned models.
- Parameter-Efficient Fine-Tuning of Large Language Models — Ding et al., EMNLP 2023. Comprehensive survey of LoRA, prefix tuning, adapters, prompt tuning, and efficiency trade-offs.
- Evaluating Large Language Models — Chang et al., EMNLP 2023. Capability benchmarks, LLM-as-judge, and evaluation pitfalls.
- Alignment for NLP: Reward Modeling, RLHF, and Beyond — Lambert, Zhu, Hajishirzi (Ai2), ACL 2024. End-to-end alignment pipeline with a focus on NLP tasks; accompanies the RLHF Book.
- LLM Agents: Architectures, Benchmarks, and Training — Yao, Narasimhan (Princeton), NAACL 2024. ReAct, tool use, CodeAct, and evaluation benchmarks for language agents.
- Preference Learning and Alignment in the LLM Era — Rafailov, Ethayarajh, Finn (Stanford), ACL 2025. DPO theory, KTO, ORPO, and on-policy preference optimization.
Active research frontiers where the field is genuinely uncertain:
- Reward hacking at scale. Persistent across techniques. RLVR mitigates but doesn't eliminate. Active work on adversarial reward shaping, judge calibration, process rewards.
- Process vs. outcome rewards. Where to invest? PRMs are dense but expensive to label and over-credit easy steps. ORMs are sparse but easier.
- Synthetic-data flywheels and model collapse. Self-improvement loops show strong gains (Self-Challenging Agents) but distribution-collapse risks remain unresolved.
- Long-horizon training. Episodes >400 turns (Verlog) push the boundary. Sample efficiency and credit assignment over hundreds of decisions is open.
- Multi-agent and user-in-the-loop RL. MUA-RL points the way; production-deployment patterns immature.
- Generalization of agent skills across environments. SWiRL's HotPotQA→GSM8K transfer is suggestive but not understood.
- Continual / just-in-time learning without gradient updates. ACE, JitRL, persona vectors/subnetworks — alternatives to retraining gaining traction.
- Open vs. closed recipes. Ai2's Tülu 3 sets an openness bar; commercial players continue to selectively close.
- Personalization without retraining. A consumer-product imperative; technically immature.
- Eval-aware models and contamination. Models recognizing eval environments (Docent finding); a structural problem for the benchmark canon.
This list is opinionated and incomplete by design. Send a PR if a tool, paper, dataset, benchmark, platform, or company belongs here.
Suggested repo hygiene:
- Run
awesome-lintbefore submitting major changes. - Prefer stable technical descriptions over market rumors.
- Include a public artifact link whenever possible.
- For company funding / valuation / ARR / headcount claims, include a verification date or avoid the claim.
- Keep one-line descriptions concise enough for scanning.
Criteria:
- Public artifact (code, paper, model, dataset, benchmark) or material commercial relevance.
- Active in the last 18 months, unless historically foundational.
- Direct relevance to post-training, fine-tuning, reward modeling, synthetic data, evaluation, serving, or agentic RL for foundation models.
For non-US contributions of equivalent rigor (DeepSeek, Qwen, Mistral, Cohere, BAAI, etc.), include — geography in the title is a starting point, not a fence.
CC0 — public domain. Take it, fork it, improve it.
Heavily indebted to Nathan Lambert and the RLHF Book (rlhfbook.com) for the conceptual architecture, to the Allen Institute for AI for the most openly documented post-training recipes in the field, and to the research and open-source community whose papers and code are the primary sources for everything here.
"In the end, with how impossible it is to measure human preferences, RLHF will never be a solved problem." — Nathan Lambert