LESSON wm.4.2 · Assessment · 150 min

Evaluate world models: from impressive to useful

Evaluate world models: from impressive to useful. Learn controllability, long-horizon consistency, physics, planning utility, sim-to-real, and safety and complete: Build an evaluat

Learning objectives

  1. Explain what problem “Evaluate world models: from impressive to useful” solves without hiding behind terminology.
  2. Trace the variables and causal links across controllability, long-horizon consistency, physics.
  3. Complete “Build an evaluation matrix across Genie, JEPA, Marble, and Cosmos” and judge the result with evidence rather than intuition.

Core concepts

controllability

controllability is an independent acceptance dimension between an impressive video and a decision-useful world model. Fix actions and initial conditions, then measure response, drift, physical constraints, planning gain, and real-world transfer separately.

long-horizon consistency

long-horizon consistency is an independent acceptance dimension between an impressive video and a decision-useful world model. Fix actions and initial conditions, then measure response, drift, physical constraints, planning gain, and real-world transfer separately.

physics

physics is an independent acceptance dimension between an impressive video and a decision-useful world model. Fix actions and initial conditions, then measure response, drift, physical constraints, planning gain, and real-world transfer separately.

planning utility

planning utility is an independent acceptance dimension between an impressive video and a decision-useful world model. Fix actions and initial conditions, then measure response, drift, physical constraints, planning gain, and real-world transfer separately.

sim-to-real

sim-to-real is an independent acceptance dimension between an impressive video and a decision-useful world model. Fix actions and initial conditions, then measure response, drift, physical constraints, planning gain, and real-world transfer separately.

Build and verify

Build an evaluation matrix across Genie, JEPA, Marble, and Cosmos

  • Predict: write the expected output, trend, or failure before running code.
  • Build: implement only the minimum components needed to answer the question.
  • Verify: compare with a baseline or trusted implementation; save seeds, parameters, and raw outputs.
  • Transfer: change one shape, dataset, scale, or workload condition and explain whether the conclusion still holds.

Open the complete interactive lesson