Skip to content

About

A curated map of foundation-model post-training: SFT, preference data, reward modeling, DPO, RLHF, RLVR, agentic RL, synthetic data, evaluation, open recipes, and production fine-tuning tools.

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

Awesome Foundation Model Fine-Tuning Awesome

A curated, opinionated catalogue of state-of-the-art tools, research, and companies for post-training and fine-tuning foundation models — covering instruction tuning, preference optimization, RLHF, RLVR, and agentic reinforcement learning. Organized around Nathan Lambert's RLHF Book (rlhfbook.com) as the canonical curriculum, then extended with industry tooling, US-university research, and the startup landscape.

Foundation-model post-training is the suite of techniques that turns a base model trained on next-token prediction into something useful: an instruction-follower, a chat assistant, a reasoner, a tool-using agent. The field has moved fast — from "SFT + simple RLHF" in 2023 to a dense ecosystem of preference optimization, verifiable rewards, process rewards, on-policy distillation, agentic RL, and synthetic-data flywheels in 2026. This list tries to be the map that wasn't there when the field exploded.

Scope. Open-source frameworks, reproducible recipes, benchmarks, datasets, key papers, and the companies (big tech, startups, non-profit labs) shaping the field. Heavy emphasis on US-based contributions but not exclusive.

Curation principle. Inclusion requires either a public artifact (code, paper, model, dataset, benchmark) or material commercial relevance. No vaporware.

Maintenance principle. Prefer durable technical descriptions over fast-decaying valuation, headcount, or revenue claims. When market details are included, add a verification date or move them to a separate landscape note.


Contents


How to Use This List

This list is intentionally structured as both a learning path and an implementation map.

  • Learning path: Foundational Reading → SFT → Preference Data → DPO → RLHF / RLVR → Agentic RL.
  • Practitioner path: Datasets → Training Frameworks → Evaluation → Serving and Multi-Adapter Infrastructure.
  • Founder / enterprise path: Managed platforms → Data operations → Evaluation, safety, and governance → Serving.
  • Research path: Open recipes → University groups → Open questions.

Suggested tags used throughout:

  • [paper] — research paper or technical report.
  • [code] — runnable implementation, framework, or library.
  • [dataset] — training, preference, reward, or evaluation dataset.
  • [benchmark] — evaluation benchmark or leaderboard.
  • [platform] — managed or commercial system relevant to fine-tuning / post-training.
  • [recipe] — reproducible training pipeline, model card, or training mixture.

Foundational Reading

The texts to read before anything else.


The Canonical Post-Training Recipe

The Lambert formulation, with chapter cross-references to The RLHF Book:

  1. Instruction Fine-Tuning (SFT) — Ch. 4. Teach the chat format and basic instruction following.
  2. Preference Data Collection — Ch. 6. Build preference pairs (or rubrics, or verifiable signals).
  3. Reward Modeling — Ch. 5/7. Train a scalar reward model on preference data; optionally outcome (ORM) or process (PRM).
  4. Rejection Sampling — Ch. 9/10. Generate many completions, keep the best by RM, SFT on them. The cheapest RLHF.
  5. Direct Alignment — Ch. 8/12. DPO and family — optimize the RL objective directly from preference pairs.
  6. Policy-Gradient RL — Ch. 6/11. PPO, GRPO, REINFORCE++. The heaviest but generally most effective stage.
  7. RLVR — Verifiable rewards for math, code, structured tasks. The 2024-2026 renaissance.
  8. Character/Product Training — Ch. 17. Synthetic-data-heavy persona shaping. The "tricks" everyone uses, nobody publishes.
  9. Evaluation — Ch. 18/19. Continuous, multi-axis, with regression gates against general capability benchmarks.

Modern recipes mix-and-match these stages; the Tülu 3 sequence (SFT → DPO → RLVR) is currently the strongest public baseline.


Stage 1 — Instruction Fine-Tuning (SFT)

Key papers

Datasets


Stage 2 — Preference Data & Reward Modeling

Key papers

Datasets

Tools

Data curation and feedback operations

  • Argilla [code] [data] — human and AI feedback workflows for preference data, evaluation sets, and dataset review.
  • Label Studio [code] [data] — general-purpose labeling and review platform that can support preference, safety, and task-specific annotation.
  • Lilac [code] [data] — dataset exploration, clustering, search, and curation for LLM training and evaluation data.
  • Cleanlab [code] [data-quality] — data-quality tooling useful for finding label issues, outliers, and low-confidence examples before fine-tuning.

Stage 3 — Direct Alignment Algorithms (DPO and family)

Key papers


Stage 4 — Reinforcement Learning (PPO, GRPO, REINFORCE)

Key algorithms and papers

Reference implementations (in The RLHF Book code library)

The book ships canonical implementations of PPO, REINFORCE, GRPO, and RLOO in code/policy_gradients/.


Stage 5 — Reinforcement Learning with Verifiable Rewards (RLVR)

The 2024-2026 renaissance. Replace the reward model with a programmatic verifier where ground truth exists.

Key papers

  • Tülu 3 / RLVR — Ai2 (Lambert et al., 2024). The paper that introduced RLVR as a named technique and scaled it. Gains of 1.7 / 3.3 / 1.3 points over DPO checkpoints on MATH / GSM8K / IFEval; scaled better at 405B than smaller scales.
  • DeepSeek-R1 — DeepSeek (2025). GRPO + RLVR for reasoning. The proof point.
  • OpenAI o1 system card — OpenAI (2024). The closed-source precursor.
  • Let's Verify Step by Step — Lightman et al. (OpenAI, 2023). PRMs for math.
  • Reinforcement Pre-Training (RPT) — multiple groups. RL applied during/before SFT, not after.
  • Golden Goose (2026) — synthesize unlimited RLVR tasks from raw internet text; 4B model trained on cyber-domain data surpasses 7B domain-specialist.

Tooling and recipes

  • open-instruct (Ai2) — the canonical RLVR codebase. Used to train Tülu 3 and OLMo 2/3.
  • verl — production-grade GRPO/PPO training, used heavily for reasoning RL.
  • VeRL-Pipeline (HybridFlow paper).

Stage 6 — Agentic / Multi-Turn / Tool-Use RL

The current frontier. Multi-step, tool-using, long-horizon agents trained with RL on trajectories, not single completions.

Key papers

  • AgentFlow / Flow-GRPO — Stanford (ICLR 2026 Oral). In-the-flow planner training: 17.2% gain with online RL vs. 19.0% collapse with SFT — strongest published evidence that SFT actively hurts on agentic tasks and RL is required.
  • Self-Challenging Agents — Meta + UC Berkeley (2025). RL on self-generated synthetic tasks: 95.8% relative improvement on Llama-3.1-8B across four tool-use environments.
  • SWiRL: Step-Wise Reinforcement Learning — Stanford + Google DeepMind (Goldie, Mirhoseini, Manning, 2025). 21.5%, 12.3%, 14.8%, 11.1%, 15.3% gains on GSM8K, HotPotQA, CofCA, MuSiQue, BeerQA; cross-task generalization (training on HotPotQA improves GSM8K by 16.9%).
  • Verlog — CMU (2025). Multi-turn RL framework for episodes >400 turns. The state-of-the-art for long-horizon training.
  • MUA-RL — Multi-turn user-agent RL with LLM-simulated users in the training loop.
  • CodeAct — UIUC + Apple (2024). Executable code as the action format with execution feedback. Foundational for code agents.
  • SWE-agent — Princeton (2024). Agent-computer interface for software engineering.
  • τ-bench — Shunyu Yao et al. (Princeton, 2024). Tool-agent-user benchmark with the pass^k reliability metric.
  • VerlTool (2025). Holistic agentic RL with comprehensive tool support — useful comparison table of frameworks.
  • Agentic Context Engineering (ACE) — Stanford + UC Berkeley + SambaNova (2025). Self-improvement without weight updates; >10pp gain on AppWorld, 8.6% on financial reasoning.

Frameworks

  • OpenHands (CMU / Neubig group). The open-source agent framework with most mindshare.
  • Verlog (CMU) — long-horizon training framework.
  • Torchforge (Meta PyTorch) — production-style agentic RL.
  • OpenEnv (Meta PyTorch) — standard environment interface for agent RL.
  • VerlTool — verl extended with comprehensive tool support.
  • NeMo-Gym (NVIDIA) — agent-based RLHF with external evaluation environments; integrates with OpenRLHF.

Benchmarks

  • SWE-bench / Verified / Multimodal / Lite (Princeton) — the de facto coding-agent benchmark. 500 verified instances, Docker-based harness, Modal-based parallel evaluation. Used by every major lab.
  • τ-bench (Sierra Research) — dynamic user-agent interaction.
  • TheAgentCompany (CMU) — multi-role enterprise agent benchmark; 3,000 hours of researcher labor to build.
  • WebArena (CMU) — realistic web agent benchmark.
  • AgentBench — multi-environment agent eval.
  • InterCode (Princeton) — interactive coding tasks.

Stage 7 — On-Policy Distillation and Model Merging

Distillation

Merging

  • MergeKit (Arcee AI) — production merging library.
  • TIES-Merging — Yadav et al. (UNC, 2023).
  • DARE — Yu et al. (Microsoft + Princeton, 2023).
  • SLERP — geometric merge; widely used baseline.
  • Evolutionary Model Merge — Sakana AI (2024). Automated search over merge configurations.

Stage 8 — Character, Personality, and Product Training

The Lambert "tricks chapter" — what frontier labs actually do to round out models. The RLHF Book Ch. 17 expanded significantly in the print edition.

Key references


Cross-Cutting — Synthetic Data Generation

Arguably the most under-credited multiplier in the modern stack.

Tools

  • distilabel (Argilla / Hugging Face) — production synthetic-data pipelines.
  • Magpie — self-aligning data generation from instruct models.
  • Curator (Bespoke Labs) — production synthetic dataset creation.
  • DataDreamer — synthetic-data generation, instruction-data workflows, and reproducible LLM data pipelines.
  • DSPy (Stanford NLP) — programmatic prompt / data / pipeline optimization; useful for generating and optimizing task-specific training traces.

Cross-Cutting — Evaluation

Capability benchmarks

Open-ended / chat evaluation

  • Chatbot Arena (LMSYS / now Arena Intelligence) — human preference Elo across 6M+ votes.
  • AlpacaEval 2.0 (Stanford). LLM-judge, length-controlled.
  • MT-Bench (LMSYS) — multi-turn LLM-judge.
  • Arena-Hard (LMSYS) — automated, hard prompts.
  • WildBench (Ai2) — real user prompts.

Agent/tool-use evaluation

Eval frameworks

  • Inspect AI (UK AI Safety Institute / now AI Security Institute). Probably the most adopted modern eval framework. Schema-compatible with Docent.
  • lm-evaluation-harness (EleutherAI). Long-standing reference for capability eval.
  • LightEval (Hugging Face).
  • OpenAI Evals.
  • HELM (Stanford CRFM). Holistic evaluation.
  • Vals AI — third-party eval platform with public leaderboards.
  • Epoch AI Benchmarks — independently reproduced frontier evals.

Safety, risk, and governance evaluation

  • NIST AI Risk Management Framework — Generative AI Profile [governance] — risk-management reference for generative AI systems.
  • MLCommons AILuminate [benchmark] [safety] — safety benchmark family for assessing model behavior across hazard categories.
  • OWASP Top 10 for LLM Applications [security] — practical taxonomy for prompt injection, data leakage, model theft, supply-chain risks, and other LLM-app threats.
  • HarmBench [benchmark] [safety] — standardized harmful-behavior evaluation for LLMs.
  • WMDP [benchmark] [safety] — hazardous-knowledge evaluation for biosecurity, cybersecurity, and chemical domains.
  • CyberSecEval [benchmark] [security] — cybersecurity-oriented model safety and capability evaluations.

Post-training regression checklist

  • General capability should not collapse after SFT / DPO / RL.
  • Instruction following should improve without overfitting to template artifacts.
  • Safety refusals should improve without excessive false refusals.
  • Reward models and LLM judges should be audited for verbosity, sycophancy, and style bias.
  • Agentic training should report trajectory-level success, pass^k reliability, tool-call validity, and failure modes.

Cross-Cutting — Interpretability and Behavior Analysis

Tools

  • Docent (Transluce, ex-Berkeley/MIT). Agent transcript analysis — clustering, rubrics, counterfactual intervention, "reward hacks across training steps" monitoring. The standard for agentic behavior analysis. GitHub.
  • Goodfire Ember — interpretability platform for model-behavior analysis, steering, training-phase introspection, and production monitoring.
  • Neuronpedia — community SAE/feature visualization.
  • SAELens — sparse-autoencoder library.
  • TransformerLens — mechanistic interpretability workbench.

Key papers


Open Models with Fully Open Post-Training Recipes

The shortlist of models that release weights + data + training code + recipes.

  • OLMo 2 / OLMo 3 (Ai2). 7B / 13B / 32B. Fully open, including Tülu-3 post-training pipeline.
  • Tülu 3 / Tülu 3 405B (Ai2). The reference open recipe. Tülu 3 SFT mixture, DPO data, RLVR setup all released.
  • Llama Instruct family (Meta). Weights open, recipe partially documented. Lambert was directly involved in some of these.
  • Qwen Instruct family (Alibaba). Strong open recipes for chat and reasoning.
  • DeepSeek-R1 / DeepSeek-V3 (DeepSeek). The RLVR-trained reasoning models. Recipe partially open.
  • Zephyr 7B (Hugging Face). The Alignment Handbook reproducible reference.
  • SmolLM3 (Hugging Face). Small-model post-training reference.
  • Nemotron families (NVIDIA). Heavily synthetic-data post-training.
  • Phi-4 (Microsoft). Synthetic-data-led small model recipe.

Training Frameworks (SFT-focused)

Framework Strength Best for
Axolotl YAML configs, broad model support, multi-GPU via FSDP2 Production multi-GPU SFT/RLHF
Unsloth Custom CUDA kernels, single-GPU speed, GRPO at 5GB VRAM Single-GPU experimentation
Torchtune (PyTorch) PyTorch-native, lean, QAT support Pure-PyTorch shops
LLaMA-Factory Broad UI/model support Quick experimentation
Hugging Face TRL Post-training library for SFT, DPO, PPO, GRPO, reward modeling, and related methods Anyone in HF ecosystem
Hugging Face PEFT LoRA, QLoRA, IA3, prefix tuning, and adapter-style parameter-efficient fine-tuning Adapter-first workflows
bitsandbytes / QLoRA 8-bit / 4-bit quantization and low-memory fine-tuning Single-GPU and cost-sensitive fine-tuning
Accelerate Lightweight distributed training launcher and device abstraction Multi-GPU without heavy infra
DeepSpeed ZeRO optimization, distributed training, memory efficiency Full fine-tuning at scale
Liger Kernel Triton kernels for memory-efficient LLM training Throughput and memory optimization
Levanter (Stanford CRFM) JAX/TPU, reproducible, named-axis tensors TPU users
NVIDIA NeMo Production scale, Megatron-LM under the hood NVIDIA-stack enterprises
ColossalAI Distributed training infra Multi-node scaling
xtuner (Shanghai AI Lab) Wide model coverage Asia-developer-heavy

Training Frameworks (RL-focused)

The 2024-2026 RL-framework explosion. From the VerlTool comparison:

Framework Origin Style Notable
verl ByteDance Seed Sync, Ray + FSDP/Megatron HybridFlow paper; widely adopted production-grade
OpenRLHF Community-led Sync/async, Ray + vLLM + ZeRO-3 Scalable PPO / DPO / GRPO / REINFORCE-style RLHF and agentic RL workflows
TRL Hugging Face Sync, single-node to small distributed Most-used learning entry point for SFT, DPO, PPO, GRPO, and reward modeling
open-instruct Ai2 Sync, Tülu/OLMo backbone The Tülu 3 / RLVR reference
AReaL Ant Research Fully async Scales to long-horizon reasoning
ROLL Alibaba Async Puzzle-environment focus
slime THUDM Async Emerging
Rlinf SAIL Singapore Async
RL2 Open Async
NeMo-Aligner + NeMo-Gym NVIDIA Sync Production with environment integration
Torchforge + OpenEnv Meta PyTorch Async Agentic-first, env standard
VerlTool TIGER-Lab Sync verl extended for tool-use
RAGEN RAGEN-AI / CMU Multi-turn Long-horizon agent training
trlX (legacy) CarperAI Sync Pre-2024 reference; NeMo-backed for large scale

Serving and Multi-Adapter Infrastructure

Critical for any commercial post-training stack — serve many fine-tuned variants from one base.

  • vLLM (Berkeley Sky Lab → vLLM project). PagedAttention; the de facto open inference engine. Native LoRA hot-swap.
  • SGLang — high-throughput inference with strong RL integration.
  • TensorRT-LLM (NVIDIA).
  • Text Generation Inference (TGI) (Hugging Face).
  • LoRAX (Predibase) — multi-LoRA serving from a single base.
  • Punica — efficient multi-tenant LoRA.
  • S-LoRA — Berkeley. Foundational paper on scalable LoRA serving.

Research Labs and University Groups

Stanford

UC Berkeley

CMU

  • Graham Neubig — OpenHands, TheAgentCompany; agent systems and evaluation.
  • Aditi Raghunathan — robustness, post-training generalization.
  • CMU Advanced NLP (Spring 2025) — uses OpenRLHF as teaching framework.
  • Verlog / RAGEN-AI — long-horizon agent training.

Princeton

MIT

Allen Institute for AI (Ai2)

  • The non-profit benchmark for open post-training. Nathan Lambert, Hanna Hajishirzi, Noah Smith, Yejin Choi, Pradeep Dasigi.
  • Tülu 3, OLMo 2/3, RewardBench, open-instruct, WildBench.

Other major US groups

  • Cornell — Alexander Rush, Claire Cardie; reasoning, agent training.
  • Harvard — Stuart Shieber; test-time interaction scaling collaborations with CMU.
  • Columbia — SWE-Bench-CL continual learning evaluation.
  • UIUC — CodeAct (Heng Ji's group).
  • UCLA — Kai-Wei Chang; agent reasoning.
  • University of Washington — Yejin Choi (now at Stanford), Hanna Hajishirzi (Ai2), Tengyang Xie (Markov-state framing of RL post-training).
  • NYU — Kyunghyun Cho, Sam Bowman (now Anthropic); preference learning theory.

Startups and Companies

Foundation-model post-training thesis bets

  • Reflection AI — Berkeley / DeepMind-rooted company centered on RL post-training for autonomous agents. A thesis-level competitor to watch.
  • Imbue (formerly Generally Intelligent). Berkeley-rooted work on robust reasoning, coding agents, and agentic systems.
  • Adept (now partially absorbed into Amazon AGI). Agent foundation models.

Fine-tuning platforms

  • Predibase — managed fine-tuning, RFT / GRPO workflows, and LoRAX multi-LoRA serving.
  • Together AI — broad fine-tuning and inference platform with support for customization workflows across open models.
  • Fireworks AI — inference and fine-tuning platform with an agent / reinforcement-fine-tuning orientation.
  • Anyscale — Ray-based ML infrastructure and managed compute for large-scale training and serving.
  • Modal — serverless GPU infrastructure; widely used for evaluation, batch inference, and ML workloads.
  • Databricks + MosaicML — enterprise data + model-training stack for fine-tuning and domain adaptation.
  • Snowflake — enterprise data platform with model customization and AI application workflows.

Managed fine-tuning / model customization platforms

Interpretability / evaluation

  • Transluce — non-profit. Docent (transcript analysis), Monitor (interpretability). Used to evaluate Claude 4, by governments for risk assessment.
  • Goodfire — interpretability platform for model behavior analysis, feature steering, and production monitoring.
  • Patronus AI — eval/safety platform.
  • Arize — observability with LLM extensions.

Agentic product companies (downstream consumers of fine-tuning)

  • Cognition AI (Devin). Acquired Windsurf in 2025. Hundreds of thousands of merged PRs.
  • Cursor (Anysphere) — AI-native coding environment and downstream consumer of code-agent advances.
  • Sierra — customer-support agents; created τ-bench as a realistic user-agent benchmark.
  • Harvey — legal AI agents and workflows.
  • Decagon — customer support.
  • Glean — enterprise search/agents.
  • Hippocratic AI — clinical agents.
  • Crescendo / Ada — CX agents.

Framework / orchestration plays

  • LangChain — agent / LLM application framework company behind LangChain, LangGraph, and LangSmith.
  • LlamaIndex — RAG / agent framework.
  • CrewAI — multi-agent orchestration.
  • Pydantic AI (Pydantic / Samuel Colvin) — typed agent framework.

Data labeling and synthetic data

  • Scale AI — preference data at scale (Meta acquired majority stake in 2025).
  • Surge AI — high-quality human feedback.
  • Argilla (now Hugging Face) — data annotation; distilabel maintainer.
  • Bespoke Labs — Curator framework.
  • Snorkel AI — programmatic data + alignment.

Inference infrastructure / serving

  • Baseten — ML deployment.
  • Replicate — model hosting.
  • Hugging Face — TRL, Alignment Handbook, datasets, Inference Endpoints, distilabel via Argilla. The hub.

Berkeley Sky Lab pipeline (institutionally)

  • Anyscale, Databricks (older AMPLab spinout), Arena Intelligence (Chatbot Arena spinout) — examples of Berkeley Sky Lab / AMPLab-style research-to-startup transfer.

Big Tech Foundation Labs

  • Anthropic — Claude family. Constitutional AI, character training, RLAIF lineage. Interpretability publishes via transformer-circuits.pub.
  • OpenAI — GPT family. RLHF popularizer; o-series RLVR.
  • Google DeepMind — Gemini family. SWiRL, Gemma open models, Gemma Scope.
  • Meta AI / FAIR — Llama family. Llama Instruct recipes, ScaleRL, Self-Challenging Agents, Torchforge.
  • Microsoft AI — Phi series; synthetic-data leadership; large investment in OpenAI.
  • NVIDIA AI — Nemotron, NeMo, NeMo-Aligner, NeMo-Gym, HelpSteer datasets, TensorRT-LLM.
  • Apple — on-device models; CodeAct co-authorship.
  • Amazon AGI / Bedrock — Nova family; Bedrock AgentCore.
  • xAI — Grok family.

Courses, Blogs, and Newsletters

Courses

Blogs / newsletters

Podcasts


Conference Tutorials

Curated tutorials from top AI/NLP/ML venues covering RLHF, alignment, post-training, and related topics. Sorted by recency within each venue.

NeurIPS

ICML

ICLR

ACL / EMNLP / NAACL


Open Questions

Active research frontiers where the field is genuinely uncertain:

  1. Reward hacking at scale. Persistent across techniques. RLVR mitigates but doesn't eliminate. Active work on adversarial reward shaping, judge calibration, process rewards.
  2. Process vs. outcome rewards. Where to invest? PRMs are dense but expensive to label and over-credit easy steps. ORMs are sparse but easier.
  3. Synthetic-data flywheels and model collapse. Self-improvement loops show strong gains (Self-Challenging Agents) but distribution-collapse risks remain unresolved.
  4. Long-horizon training. Episodes >400 turns (Verlog) push the boundary. Sample efficiency and credit assignment over hundreds of decisions is open.
  5. Multi-agent and user-in-the-loop RL. MUA-RL points the way; production-deployment patterns immature.
  6. Generalization of agent skills across environments. SWiRL's HotPotQA→GSM8K transfer is suggestive but not understood.
  7. Continual / just-in-time learning without gradient updates. ACE, JitRL, persona vectors/subnetworks — alternatives to retraining gaining traction.
  8. Open vs. closed recipes. Ai2's Tülu 3 sets an openness bar; commercial players continue to selectively close.
  9. Personalization without retraining. A consumer-product imperative; technically immature.
  10. Eval-aware models and contamination. Models recognizing eval environments (Docent finding); a structural problem for the benchmark canon.

Contributing

This list is opinionated and incomplete by design. Send a PR if a tool, paper, dataset, benchmark, platform, or company belongs here.

Suggested repo hygiene:

  • Run awesome-lint before submitting major changes.
  • Prefer stable technical descriptions over market rumors.
  • Include a public artifact link whenever possible.
  • For company funding / valuation / ARR / headcount claims, include a verification date or avoid the claim.
  • Keep one-line descriptions concise enough for scanning.

Criteria:

  • Public artifact (code, paper, model, dataset, benchmark) or material commercial relevance.
  • Active in the last 18 months, unless historically foundational.
  • Direct relevance to post-training, fine-tuning, reward modeling, synthetic data, evaluation, serving, or agentic RL for foundation models.

For non-US contributions of equivalent rigor (DeepSeek, Qwen, Mistral, Cohere, BAAI, etc.), include — geography in the title is a starting point, not a fence.


License

CC0 — public domain. Take it, fork it, improve it.

Acknowledgments

Heavily indebted to Nathan Lambert and the RLHF Book (rlhfbook.com) for the conceptual architecture, to the Allen Institute for AI for the most openly documented post-training recipes in the field, and to the research and open-source community whose papers and code are the primary sources for everything here.

"In the end, with how impossible it is to measure human preferences, RLHF will never be a solved problem." — Nathan Lambert

About

A curated map of foundation-model post-training: SFT, preference data, reward modeling, DPO, RLHF, RLVR, agentic RL, synthetic data, evaluation, open recipes, and production fine-tuning tools.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors