LESSON 5.8 · Engineering · 110 min

Safety, red teaming, and system cards

Safety, red teaming, and system cards. Learn threat models, jailbreaks, misuse, and risk registers and complete: Red-team a small model. Part of the “Make the Model an Assistant” l

Learning objectives

  1. Explain what problem “Safety, red teaming, and system cards” solves without hiding behind terminology.
  2. Trace the variables and causal links across threat models, jailbreaks, misuse.
  3. Complete “Red-team a small model” and judge the result with evidence rather than intuition.

Core concepts

threat models

threat models is part of the lesson’s causal model. State its inputs, outputs, invariants, and failure mode; then verify it with a hand-check or a minimal experiment before moving to an optimized implementation.

jailbreaks

jailbreaks is part of the lesson’s causal model. State its inputs, outputs, invariants, and failure mode; then verify it with a hand-check or a minimal experiment.

misuse

misuse is part of the lesson’s causal model. State its inputs, outputs, invariants, and failure mode; then verify it with a hand-check or a minimal experiment.

and risk registers

and risk registers is part of the lesson’s causal model. State its inputs, outputs, invariants, and failure mode; then verify it with a hand-check or a minimal experiment.

Build and verify

Red-team a small model

  • Predict: write the expected output, trend, or failure before running code.
  • Build: implement only the minimum components needed to answer the question.
  • Verify: compare with a baseline or trusted implementation; save seeds, parameters, and raw outputs.
  • Transfer: change one shape, dataset, scale, or workload condition and explain whether the conclusion still holds.

Open the complete interactive lesson