
Every time a chatbot asks which of two answers you prefer, that click eventually becomes a gradient step. Here is the full mechanism — preference pairs, Bradley–Terry reward models, KL penalties, and policy gradients — rebuilt from scratch with four replies and honest arithmetic.
You have probably done this yourself: a chatbot shows you two draft answers and asks which one you prefer. It feels like a survey. It is actually the most consequential training signal in modern AI — the raw ore of reinforcement learning from human feedback, the technique that turned a text predictor into an assistant that follows instructions. The method was codified in the InstructGPT paper in 2022 , and its roots go back to 2017 work on learning from human preferences
.
But how does a preference become math? Most explanations wave at three boxes labeled SFT, reward model, and PPO, and call it a day. This post does the opposite: we build all of RLHF from scratch on a toy chatbot with only four possible replies, with every probability and gradient visible — and then run the full experiment end to end, 48,000 sampled replies, to watch the numbers actually converge.
A real language model can write endlessly different answers. That is what makes RLHF hard to see. So shrink the problem: a chatbot that answers one question — how do I reverse a list in Python? — and can only produce four replies: the right one, a chatty one, a wrong one, and a refusal. It starts out giving them with probabilities 0.20, 0.40, 0.30, and 0.10.
In place of billions of parameters, the model has four numbers — one per reply — passed through a softmax: exponentiate each, divide by the sum. Push one number up and its probability rises while the others shrink. Every mechanism in RLHF can be demonstrated on these four numbers, and nothing is hidden. The goal: make right fire more often than 0.20.
The standard first move is supervised fine-tuning (SFT): a person writes the ideal reply, and training minimizes the negative log-probability of that written answer. The gradient is delightfully simple — one-hot minus probabilities — so the first update pushes right up by 0.8 and everything else down. Five steps take right from 0.20 to 0.84. Solved?
Not quite. SFT has four structural problems, and they are worth internalizing because they explain why the entire RLHF apparatus exists:
Reinforcement learning inverts the arrangement. The model writes its own answers, a reward model grades each one, and the policy is pushed toward the better ones. Learning from the current model's own outputs — on-policy learning — means the model gets feedback on the mistakes it actually makes, needs nobody to write answers, and can discover replies better than any it was shown. All it needs is a score. Which raises the real question: where do scores come from?
Ask people to write the best reply and the work is hard. Ask them to rate replies on a ten-point scale and the numbers are inconsistent — one rater's 7 is another's 4. But show a person two replies side by side and ask which is better? — that is easy, fast, and surprisingly reliable. RLHF is built on this asymmetry: comparison is cheap; absolute judgment is not.
Each comparison becomes one training triple: the prompt, the preferred reply, and the rejected reply. Our toy dataset has 1,000 such pairs — and they are noisy. The right reply beat the chatty one 198 times, but lost to it 35 times. Wrong versus refusal came out 43 to 34. No single preference is the truth; the truth, if it exists, is statistical. This is exactly the situation chess federations face with players, and they solved it decades ago.
An Elo rating compresses a player's entire win-loss record into one number, and the gap between two ratings predicts how often one beats the other. The reward model does the same thing for replies, in a form called the Bradley–Terry model: every reply gets a scalar score, and the probability that one reply is preferred over another is the sigmoid of the score gap. A gap of zero is a coin flip. A gap of 1.5 means the better reply wins about 82% of the time.

Fitting the scores is maximum likelihood, and the arithmetic is worth seeing once. Start with all scores at zero: every pair is predicted at 0.5, and the total log-likelihood of the dataset is 1,000 × log(0.5) = −693. Now spread the scores out — right well above chatty, wrong and refusal below — and the log-likelihood climbs to −478. Invert the ranking entirely (refusal on top) and it craters to −1,482. The loss is just the negated sum, so a good reward model is any set of scores that makes the observed preferences look unsurprising. Each gradient step nudges every winner up and every loser down, weighted by how wrong the current gap is — small, uncertain gaps push hardest. Three hundred steps and the scores settle.
Two elegant properties fall out. Only gaps matter, so you can pin the best reply at 2.0 and express everything relative to it. And because the model is fit on noisy data rather than a gold answer key, disagreement among raters becomes a feature: it shapes how confident the scores are.
Here is the entire architecture — the supervised starting point, the offline reward model, and the online reinforcement learning loop that uses it:
Scale this up and surprisingly little changes. In the InstructGPT pipeline, the reward model is another large language model: take the fine-tuned model, remove the final layer that scores the next token, and train it to emit a single number for a prompt–reply pair. Its training loss is exactly the toy loss — the negated log-sigmoid of the preferred reply's score minus the rejected one's, averaged over human comparisons . The Llama 3 generation of open models trains its reward model the same way on top of its pre-trained checkpoint, and for some prompts adds a third, human-edited reply, so each ranking reads edited beats chosen beats rejected
. The toy's four numbers become billions, but the object being fit is identical.
Now the loop. Each round, the chatbot samples a batch of fresh replies. Each reply earns a reward — the reward model's score — minus a penalty: β times the log-ratio of the new model's probability to the frozen reference's. Averaged over samples, that penalty is β times the KL divergence — a measurable distance between what the model is becoming and what it was.
Why the leash? Because the reward model is only trustworthy near the replies humans actually judged. Far from that territory its scores are extrapolations, and the InstructGPT paper says so plainly. Left unconstrained, a policy gradient will march the chatbot off into regions where the reward model is confidently wrong — and pile probability onto its mistakes. That failure mode has a name every practitioner learns early: reward hacking.
The update itself is REINFORCE with one standard refinement: each reward is compared against the round's average, and only the difference — the advantage — scales the push. Replies better than average get more likely; worse ones get suppressed. The baseline changes nothing in expectation but cuts the noise dramatically. This group-mean baseline is precisely the trick at the heart of GRPO, the algorithm introduced with DeepSeekMath ; InstructGPT instead used PPO, which learns a separate critic for the baseline and clips how far each update may move the probabilities
.
Reading arithmetic is one thing; running it is another. So I implemented the toy experiment in full — fitted reward scores (right pinned at 2.0), β = 0.5, learning rate 0.05, batches of 16 fresh replies, policy gradient with the group-mean baseline and the KL penalty — and let it run 3,000 rounds: 48,000 sampled replies in all.

The result: 0.911 for the right reply, at a final KL of 1.21 from the reference model. Two things are worth noticing in the curve. First, the learning is smooth but not free — the KL penalty visibly taxes every step, which is why the model converges near perfect rather than slamming to 1.0. Second, the wrong reply and the chatty reply both die, but at different rates shaped by their rewards, not their initial probabilities — the exact discrimination SFT was blind to.
RLHF's three-stage machinery dominated alignment from 2022 onward, but the field has not stood still. Direct Preference Optimization showed that the preference loss can be applied to the language model directly, skipping the explicit reward model and the RL loop entirely — your language model is, in the paper's phrase, secretly a reward model . And for domains with checkable answers — math, code — reinforcement learning from verifiable rewards replaces human raters with compilers and test suites, powering the recent wave of reasoning models. The through-line is unchanged: scalar feedback, on-policy learning, and a leash that keeps the model where its feedback can be trusted.
The deepest lesson of the toy board is not any one equation. It is the division of labor: SFT teaches the model to imitate, the reward model compresses messy human judgment into a number, and the RL stage lets the model outrun its teachers — carefully, on a KL leash. Four numbers, one sigmoid, and three thousand rounds are enough to see the whole shape of it. The next time a chatbot asks which answer you prefer, you will know exactly what your click is about to become.
| Supervised fine-tuning | RLHF | |
|---|---|---|
| Feedback | A human-written ideal answer | A pairwise preference (cheap, noisy) |
| Data | Fixed demonstrations | Fresh, on-policy samples each round |
| Signal | −log p of the written reply | Advantage: reward − batch mean, with KL penalty |
| Knows which wrong is worse | No — pushes down by probability | Yes — the reward model ranks everything |
| Ceiling | A perfect copy of its writers | Bounded by reward-model quality, not by demonstrations |
Methods and history as published in the original papers: InstructGPT , Christiano et al. 2017
, PPO
, DeepSeekMath / GRPO
, DPO
, Llama 3
.

Thoughts and essays, published with Yokush. See more posts
Comments 0