ExplainerHow a Single Click Trains a Chatbot: RLHF From First Principles
RLHF turns cheap pairwise human preferences into a trained reward model via Bradley–Terry, then uses policy gradients with a KL leash to push the chatbot toward better replies without reward hacking. …