🎯 RLHF & Alignment
Aligning LLMs with human preferences using RLHF, DPO, and modern preference optimization methods
On this page
Before you start
Read fine-tuning for supervised post-training, RL basics for policy gradients and PPO, and information theory for log probabilities and KL divergence. You should be comfortable distinguishing a sampled reward from an expected objective.
After this chapter, you should be able to turn a response preference into a training loss, distinguish the policy, reference, reward, and critic roles, derive the DPO log-ratio margin, and design evaluation that detects preference shortcuts and capability regressions.
The problem: two plausible answers, one preferable
A coding assistant can produce syntactically valid answers that differ in correctness, clarity, and security. Supervised fine-tuning needs a target response. Preference learning instead asks which of two responses better satisfies an explicit rubric. That supplies comparative evidence, but does not establish an absolute quality score or make the annotator infallible.
A dataset entry contains a prompt x, a preferred response y_w, and a rejected response y_l. The preference might come from a person, a judge model, executable evidence, or a combination. Record the source and rubric. Correctness, brevity, politeness, and refusal behavior can conflict; silently pooling inconsistent rubrics teaches an ambiguous objective.
Worked example: a DPO update from one pair
Suppose the preferred answer includes a necessary bounds check and the rejected answer omits it. The model and a frozen reference assign these sequence log probabilities, using natural logarithms:
| Response | Current log probability | Reference log probability | Current minus reference |
|---|---|---|---|
| Preferred y_w | −2.0 | −3.0 | +1.0 |
| Rejected y_l | −3.0 | −2.5 | −0.5 |
The preferred response has improved relative to the reference by more than the rejected response. The relative margin is m = 1.0 − (−0.5) = 1.5. At β = 0.2:
z = βm = 0.3
σ(z) = 1 / (1 + exp(−0.3)) ≈ 0.5744
loss = −log σ(z) ≈ 0.5544
At initialization equal to the reference, m = 0 and the loss is log 2 ≈ 0.6931. Reducing this pair's loss increases its relative preference margin. It does not guarantee that the preferred response's absolute probability rises: both response probabilities could fall while the rejected one's falls more.
Notation and preference models
| Symbol | Meaning |
|---|---|
| x, y_w, y_l | Prompt, preferred response, rejected response |
| πθ, πref | Trainable policy and fixed reference distribution |
| rφ(x,y) | A learned scalar reward score |
| σ(z) | Logistic function, 1/(1 + exp(−z)) |
| β > 0 | KL regularization coefficient in the convention derived below |
| DKL(π ‖ πref) | Expected log probability ratio under π |
A Bradley–Terry preference model assumes P(y_w preferred to y_l given x) = σ(rφ(x,y_w) − rφ(x,y_l)). Its loss is the negative log probability of the observed preference. Only reward differences are identified: adding a prompt-dependent constant to both scores changes no prediction. A deterministic preference label still represents uncertain human judgment; disagreement, ties, and ambiguous pairs deserve explicit handling.
The InstructGPT-style RLHF pipeline
- Supervised fine-tuning: train on demonstrations to establish desired response behavior. SFT is a useful starting point, not a mathematical requirement for every RL experiment.
- Reward modeling: collect comparisons and fit a reward model. Evaluate its ranking accuracy and failure modes on held-out prompts and newly generated responses.
- Policy optimization: generate fresh responses and optimize reward with reference regularization, often using PPO and a learned value critic.
For each prompt, the idealized objective is:
J(π) = E_y~π[r(x,y)] − β DKL(π(· given x) ‖ πref(· given x)).
The policy produces responses. The reference supplies baseline probabilities. The reward model scores completed outputs. The value critic estimates future return from partial prefixes for advantage estimation. These are four roles, not a rule that four full models must reside on one GPU: implementations may share a backbone, use adapters, cache fixed scores, shard, or offload. Optimizer and activation memory still need separate accounting.
PPO clipping limits a sampled surrogate's incentive in one direction. It does not enforce a hard trust region. A KL penalty discourages drift from the reference but cannot guarantee coherence, prevent reward hacking, or enforce an application permission boundary.
Deriving DPO from the regularized optimum
For a fixed reward and reference with adequate support, optimizing the idealized objective gives:
π*(y given x) = πref(y given x) exp(r(x,y)/β) / Z(x).
Here Z(x) normalizes the probabilities. Rearranging:
r(x,y) = β[log π*(y given x) − log πref(y given x)] + β log Z(x).
The prompt-dependent normalizer cancels when taking the reward difference of two responses. DPO parameterizes that difference using the trainable policy and fits preferences directly:
m = [log πθ(y_w given x) − log πref(y_w given x)]
− [log πθ(y_l given x) − log πref(y_l given x)]
z = βm
L_DPO = −mean over preference pairs [log σ(z)]
This avoids separately fitting a reward model and generating rollouts during the classic offline DPO update. It is not a claim that finite-data DPO training and a particular PPO run produce identical policies. The derivation assumes the reward/preference model and a KL-regularized optimum; data coverage, model capacity, optimization, and annotation errors remain.
The reference matters: it changes the relative margin and the implied regularization. At the exact fixed-reward optimum, larger β weakens reward-driven deviation. In finite DPO training, varying β also changes logits and gradients; measured policy KL need not change monotonically. DPO is not a hard KL constraint or a guaranteed trust region. Measure drift and downstream behavior rather than treating β as a promise.
Implementing the objective correctly
For an autoregressive model, log π(y given x) is the sum of response-token log probabilities conditioned on the prompt and earlier response tokens. Use consistent response masks, EOS inclusion, tokenizer, chat template, and truncation for chosen and rejected sequences. Averaging by response length instead of summing changes the objective; it is not an interchangeable numerical trick.
A practical DPO step computes policy log probabilities for both responses, subtracts frozen-reference log probabilities, constructs z, and uses a stable −logsigmoid(z) loss. Reference probabilities may be precomputed if the reference and tokenized data remain fixed. Keep reference computations detached from training. Handle empty/truncated targets and avoid probability-space products that underflow on long sequences.
The reference must assign support to the responses whose ratios you use. A mismatched tokenizer or prompt rendering can make an apparently valid numerical ratio compare different events. Track preferred/rejected log probabilities separately, margin distributions, response length, held-out preference accuracy, and generation quality.
Data and evaluation determine what is learned
Split by prompt/source family before collecting or iterating on labels. Randomize answer order to measure position bias. Include concise correct answers, verbose wrong answers, justified disagreement, safe assistance, and appropriate refusals. Code tests provide useful evidence but two responses with the same test result can still differ in maintainability, efficiency, or untested correctness; do not discard every tied-test pair automatically.
A reward rising while independent task quality falls indicates over-optimization or evaluation mismatch. Test generated responses from the updated policy, since held-out static pairs may not cover new behaviors. Human and model judges can both exhibit length, style, and knowledge biases. Calibrate judges against reviewed examples and report uncertainty and subgroup results.
Measure capability changes before and after post-training with identical decoding and grading. A lower benchmark score may reflect missing capability, over-refusal, output-format changes, or evaluator error; these require different fixes. Safety and helpfulness can trade off locally, but better labels and rubrics can improve both.
Optional alternatives and their boundaries
KTO uses unpaired desirable/undesirable feedback with a different objective. It is not a universal prescription for positive-only demonstrations. ORPO combines a supervised likelihood term with an odds-ratio preference term without a separate reference. IPO uses a finite-target squared objective; SimPO uses length-normalized policy scores and a margin. These are distinct objectives, not identical DPO implementations with components deleted.
GRPO is an online policy optimization method using group-relative rewards instead of a learned value critic. It can use learned reward models or verifiers; it is not reward-free. RLVR specifies verifiable rewards and can follow SFT. See RL basics and the GRPO playground.
Advanced variants also change token/sequence weighting, clipping, and sampling. Equal averaging of per-sequence mean losses gives each response equal weight; averaging across all tokens gives long responses more total weight. For lengths 10 and 100, token averaging assigns 100/110 of that batch's token weight to the longer response. Neither convention inherently prevents all length biases. Skipping equal-reward groups changes the sampled training distribution and should be evaluated as such.
Constitutional AI uses written principles in a supervised critique/revision phase, then uses AI-generated comparisons to train a preference reward for RL. Humans still choose the principles and evaluate outcomes. Explicit principles improve auditability of the intended rubric, not proof that every output follows it. AI feedback can reproduce or amplify a judge's blind spots.
Check your understanding
A new pair has preferred current/reference log probabilities −4 and −3, and rejected current/reference log probabilities −6 and −4. With β = 0.5, compute the margin and DPO loss. Has the preferred response become more probable than in the reference?
Solution
The preferred log ratio is −1 and the rejected log ratio is −2, so m = 1. Then z = 0.5 and the loss is −log σ(0.5) ≈ 0.4741. The preferred response is actually less probable than in the reference because its log ratio is negative. DPO fits a relative preference; lower pair loss does not guarantee higher absolute preferred-response likelihood.
Continue learning
Safety and alignment develops threat models and refusal metrics. Test-time compute explains verifiers, process supervision, and the documented R1/R1-Zero training distinction. Agentic AI covers authority and execution controls beyond model training.
References
- Ouyang et al., InstructGPT: demonstrations, comparison rewards, and policy optimization.
- Rafailov et al., DPO: KL-regularized optimum and direct preference loss.
- Shao et al., DeepSeekMath: group-relative optimization with reward signals.
- Bai et al., Constitutional AI: supervised critique/revision and AI preference feedback.