LESSON 5.5 · Theory · 150 min

PPO, GRPO, and RL post-training

PPO, GRPO, and RL post-training. Learn policy gradients, advantage, KL control, and rollouts and complete: Compare SFT, DPO, and RL on a toy task. Part of the “Make the Model an As

Learning objectives

  1. Explain what problem “PPO, GRPO, and RL post-training” solves without hiding behind terminology.
  2. Trace the variables and causal links across policy gradients, advantage, KL control.
  3. Complete “Compare SFT, DPO, and RL on a toy task” and judge the result with evidence rather than intuition.

Core concepts

policy gradients

policy gradients is part of the lesson’s causal model. State its inputs, outputs, invariants, and failure mode; then verify it with a hand-check or a minimal experiment before moving to an optimized implementation.

advantage

advantage is part of the lesson’s causal model. State its inputs, outputs, invariants, and failure mode; then verify it with a hand-check or a minimal experiment.

KL control

KL control is part of the lesson’s causal model. State its inputs, outputs, invariants, and failure mode; then verify it with a hand-check or a minimal experiment.

and rollouts

and rollouts is part of the lesson’s causal model. State its inputs, outputs, invariants, and failure mode; then verify it with a hand-check or a minimal experiment.

Build and verify

Compare SFT, DPO, and RL on a toy task

  • Predict: write the expected output, trend, or failure before running code.
  • Build: implement only the minimum components needed to answer the question.
  • Verify: compare with a baseline or trusted implementation; save seeds, parameters, and raw outputs.
  • Transfer: change one shape, dataset, scale, or workload condition and explain whether the conclusion still holds.

Open the complete interactive lesson