SiddharthAll posts
How a Single Click Trains a Chatbot: RLHF From First Principles
Kimbho Thoughts|Explainer

How a Single Click Trains a Chatbot: RLHF From First Principles

What you’ll learn
  • RLHF turns cheap pairwise human preferences into a trained reward model via Bradley–Terry, then uses policy gradients with a KL leash to push the chatbot toward better replies without reward hacking.
  • The core takeaway: SFT can only imitate its writers, while RLHF lets the model outrun them—bounded by reward-model quality, not demonstrations.

Every time a chatbot asks which of two answers you prefer, that click eventually becomes a gradient step. Here is the full mechanism — preference pairs, Bradley–Terry reward models, KL penalties, and policy gradients — rebuilt from scratch with four replies and honest arithmetic.

You have probably done this yourself: a chatbot shows you two draft answers and asks which one you prefer. It feels like a survey. It is actually the most consequential training signal in modern AI — the raw ore of reinforcement learning from human feedback, the technique that turned a text predictor into an assistant that follows instructions. The method was codified in the InstructGPT paper in 2022 arxiv.org, and its roots go back to 2017 work on learning from human preferences arxiv.org.

But how does a preference become math? Most explanations wave at three boxes labeled SFT, reward model, and PPO, and call it a day. This post does the opposite: we build all of RLHF from scratch on a toy chatbot with only four possible replies, with every probability and gradient visible — and then run the full experiment end to end, 48,000 sampled replies, to watch the numbers actually converge.

Start with the smallest honest chatbot

A real language model can write endlessly different answers. That is what makes RLHF hard to see. So shrink the problem: a chatbot that answers one question — how do I reverse a list in Python? — and can only produce four replies: the right one, a chatty one, a wrong one, and a refusal. It starts out giving them with probabilities 0.20, 0.40, 0.30, and 0.10.

In place of billions of parameters, the model has four numbers — one per reply — passed through a softmax: exponentiate each, divide by the sum. Push one number up and its probability rises while the others shrink. Every mechanism in RLHF can be demonstrated on these four numbers, and nothing is hidden. The goal: make right fire more often than 0.20.

The seduction — and the ceiling — of supervised fine-tuning

The standard first move is supervised fine-tuning (SFT): a person writes the ideal reply, and training minimizes the negative log-probability of that written answer. The gradient is delightfully simple — one-hot minus probabilities — so the first update pushes right up by 0.8 and everything else down. Five steps take right from 0.20 to 0.84. Solved?

Not quite. SFT has four structural problems, and they are worth internalizing because they explain why the entire RLHF apparatus exists:

  • Writing is expensive. A human must author a good reply for every training prompt — slow, skilled work that does not scale.
  • Many prompts have no single right answer. A birthday message can be warm, funny, or short. SFT pushes up whichever one was written and suppresses equally good alternatives.
  • SFT never learns which wrong answer is worse. In the toy step above, the chatty reply got pushed down harder than the flatly wrong one — not because it is worse, but because it was more probable. The loss is blind to degrees of badness.
  • SFT can only copy. Its best possible outcome is a perfect imitation of whoever wrote its answers.

Reinforcement learning inverts the arrangement. The model writes its own answers, a reward model grades each one, and the policy is pushed toward the better ones. Learning from the current model's own outputs — on-policy learning — means the model gets feedback on the mistakes it actually makes, needs nobody to write answers, and can discover replies better than any it was shown. All it needs is a score. Which raises the real question: where do scores come from?

Humans can't score, but they can choose

Ask people to write the best reply and the work is hard. Ask them to rate replies on a ten-point scale and the numbers are inconsistent — one rater's 7 is another's 4. But show a person two replies side by side and ask which is better? — that is easy, fast, and surprisingly reliable. RLHF is built on this asymmetry: comparison is cheap; absolute judgment is not.

Each comparison becomes one training triple: the prompt, the preferred reply, and the rejected reply. Our toy dataset has 1,000 such pairs — and they are noisy. The right reply beat the chatty one 198 times, but lost to it 35 times. Wrong versus refusal came out 43 to 34. No single preference is the truth; the truth, if it exists, is statistical. This is exactly the situation chess federations face with players, and they solved it decades ago.

From chess ratings to reward models

An Elo rating compresses a player's entire win-loss record into one number, and the gap between two ratings predicts how often one beats the other. The reward model does the same thing for replies, in a form called the Bradley–Terry model: every reply gets a scalar score, and the probability that one reply is preferred over another is the sigmoid of the score gap. A gap of zero is a coin flip. A gap of 1.5 means the better reply wins about 82% of the time.

Bradley-Terry sigmoid curve: preference probability as a function of the score gap between two replies
Fig 1 — The Bradley–Terry link function. Scores live on an unbounded scale; the sigmoid squashes their gap into a probability.

Fitting the scores is maximum likelihood, and the arithmetic is worth seeing once. Start with all scores at zero: every pair is predicted at 0.5, and the total log-likelihood of the dataset is 1,000 × log(0.5) = −693. Now spread the scores out — right well above chatty, wrong and refusal below — and the log-likelihood climbs to −478. Invert the ranking entirely (refusal on top) and it craters to −1,482. The loss is just the negated sum, so a good reward model is any set of scores that makes the observed preferences look unsurprising. Each gradient step nudges every winner up and every loser down, weighted by how wrong the current gap is — small, uncertain gaps push hardest. Three hundred steps and the scores settle.

Two elegant properties fall out. Only gaps matter, so you can pin the best reply at 2.0 and express everything relative to it. And because the model is fit on noisy data rather than a gold answer key, disagreement among raters becomes a feature: it shapes how confident the scores are.

The full pipeline on one board

Here is the entire architecture — the supervised starting point, the offline reward model, and the online reinforcement learning loop that uses it:

STAGE 0 · SUPERVISED FINE-TUNING STAGE 1 · REWARD MODEL (OFFLINE) STAGE 2 · ONLINE RL (POLICY GRADIENT) People write ideal answers Model learns to copy them Freeze as reference πref the SFT checkpoint, kept frozen Chatbot drafts reply pairs People pick the better reply 1,000 preference pairs — noisy Fit scores by max likelihood Bradley–Terry · one score per reply Sample 16 fresh replies Score: r(y) − β·log(π/πref) reward minus KL penalty Compute the advantage score − batch mean (GRPO-style) Update π toward advantage REINFORCE · PPO · GRPO scalar reward init policy + KL anchor
Fig 2 — RLHF end to end. Stage 1 runs once, offline; stage 2 loops for thousands of rounds on replies the model writes as it learns.

Inside a real reward model

Scale this up and surprisingly little changes. In the InstructGPT pipeline, the reward model is another large language model: take the fine-tuned model, remove the final layer that scores the next token, and train it to emit a single number for a prompt–reply pair. Its training loss is exactly the toy loss — the negated log-sigmoid of the preferred reply's score minus the rejected one's, averaged over human comparisons arxiv.org. The Llama 3 generation of open models trains its reward model the same way on top of its pre-trained checkpoint, and for some prompts adds a third, human-edited reply, so each ranking reads edited beats chosen beats rejected arxiv.org. The toy's four numbers become billions, but the object being fit is identical.

Stage two: learning online without drifting off a cliff

Now the loop. Each round, the chatbot samples a batch of fresh replies. Each reply earns a reward — the reward model's score — minus a penalty: β times the log-ratio of the new model's probability to the frozen reference's. Averaged over samples, that penalty is β times the KL divergence — a measurable distance between what the model is becoming and what it was.

Why the leash? Because the reward model is only trustworthy near the replies humans actually judged. Far from that territory its scores are extrapolations, and the InstructGPT paper says so plainly. Left unconstrained, a policy gradient will march the chatbot off into regions where the reward model is confidently wrong — and pile probability onto its mistakes. That failure mode has a name every practitioner learns early: reward hacking.

The update itself is REINFORCE with one standard refinement: each reward is compared against the round's average, and only the difference — the advantage — scales the push. Replies better than average get more likely; worse ones get suppressed. The baseline changes nothing in expectation but cuts the noise dramatically. This group-mean baseline is precisely the trick at the heart of GRPO, the algorithm introduced with DeepSeekMath arxiv.org; InstructGPT instead used PPO, which learns a separate critic for the baseline and clips how far each update may move the probabilities arxiv.org.

Run it. The numbers converge.

Reading arithmetic is one thing; running it is another. So I implemented the toy experiment in full — fitted reward scores (right pinned at 2.0), β = 0.5, learning rate 0.05, batches of 16 fresh replies, policy gradient with the group-mean baseline and the KL penalty — and let it run 3,000 rounds: 48,000 sampled replies in all.

Training curve: probability of each of the four replies over 3,000 rounds of policy gradient; the right answer climbs from 0.20 to 0.911 while chatty, wrong and refusal decay toward zero
Fig 3 — The full run. The right answer rises from 0.20 to 0.911; chatty, wrong, and refusal decay away. Final KL from the reference: 1.21.

The result: 0.911 for the right reply, at a final KL of 1.21 from the reference model. Two things are worth noticing in the curve. First, the learning is smooth but not free — the KL penalty visibly taxes every step, which is why the model converges near perfect rather than slamming to 1.0. Second, the wrong reply and the chatty reply both die, but at different rates shaped by their rewards, not their initial probabilities — the exact discrimination SFT was blind to.

Where the field went next

RLHF's three-stage machinery dominated alignment from 2022 onward, but the field has not stood still. Direct Preference Optimization showed that the preference loss can be applied to the language model directly, skipping the explicit reward model and the RL loop entirely — your language model is, in the paper's phrase, secretly a reward model arxiv.org. And for domains with checkable answers — math, code — reinforcement learning from verifiable rewards replaces human raters with compilers and test suites, powering the recent wave of reasoning models. The through-line is unchanged: scalar feedback, on-policy learning, and a leash that keeps the model where its feedback can be trusted.


The deepest lesson of the toy board is not any one equation. It is the division of labor: SFT teaches the model to imitate, the reward model compresses messy human judgment into a number, and the RL stage lets the model outrun its teachers — carefully, on a KL leash. Four numbers, one sigmoid, and three thousand rounds are enough to see the whole shape of it. The next time a chatbot asks which answer you prefer, you will know exactly what your click is about to become.

Supervised fine-tuningRLHF
FeedbackA human-written ideal answerA pairwise preference (cheap, noisy)
DataFixed demonstrationsFresh, on-policy samples each round
Signal−log p of the written replyAdvantage: reward − batch mean, with KL penalty
Knows which wrong is worseNo — pushes down by probabilityYes — the reward model ranks everything
CeilingA perfect copy of its writersBounded by reward-model quality, not by demonstrations

Methods and history as published in the original papers: InstructGPT arxiv.org, Christiano et al. 2017 arxiv.org, PPO arxiv.org, DeepSeekMath / GRPO arxiv.org, DPO arxiv.org, Llama 3 arxiv.org.

Image credits

Cover illustration
Generated for this article
AI-generated
0 comments
Next up

Everything Is Commodity

The core thesis is that commoditization is sweeping through all three layers of the AI stack at once—software, models, and silicon—with each layer's collapse driven partl…
Read this next
Siddharth
Siddharth

Thoughts and essays, published with Yokush. See more posts

Comments 0

Name & email required. Your email is never shown publicly.
No comments yet — be the first.