Pretraining: next-token prediction on trillions of tokens. Produces broad knowledge and skills, but a base model that continues text rather than following instructions. It is by far the most expensive stage.
Supervised fine-tuning (SFT): train on thousands to millions of curated prompt→response conversations. The model learns the assistant persona and format. RLHF trains a reward model on human preference comparisons and optimises the LLM against it (classically with PPO). DPO skips the reward model and optimises directly on preference pairs: simpler and now common.
Newer stages use reinforcement learning with verifiable rewards (unit tests passing, maths answers matching) to train long reasoning. As an engineer you mostly consume these models, but knowing the stages explains behaviour: sycophancy, refusals and verbosity are post-training artefacts; knowledge gaps are pretraining artefacts.
Going deeper
RLHF in three steps: collect human comparisons of model outputs, train a reward model to predict preferences, then optimise the LLM against the reward with a KL penalty that keeps it close to the SFT model. DPO shows the same objective can be optimised directly on preference pairs without a separate reward model.
Constitutional AI replaces many human labels with AI feedback guided by written principles, the basis of RLAIF.
Best resources for this lesson
Where this comes back
- Week 33Your own fine-tunes are small-scale SFT, usually with LoRA.