LESSON wm.4.2 · Assessment · 150 min
Evaluate world models: from impressive to useful
Evaluate world models: from impressive to useful. Learn controllability, long-horizon consistency, physics, planning utility, sim-to-real, and safety and complete: Build an evaluat
Learning objectives
- Explain what problem “Evaluate world models: from impressive to useful” solves without hiding behind terminology.
- Trace the variables and causal links across controllability, long-horizon consistency, physics.
- Complete “Build an evaluation matrix across Genie, JEPA, Marble, and Cosmos” and judge the result with evidence rather than intuition.
Core concepts
controllability
controllability is an independent acceptance dimension between an impressive video and a decision-useful world model. Fix actions and initial conditions, then measure response, drift, physical constraints, planning gain, and real-world transfer separately.
long-horizon consistency
long-horizon consistency is an independent acceptance dimension between an impressive video and a decision-useful world model. Fix actions and initial conditions, then measure response, drift, physical constraints, planning gain, and real-world transfer separately.
physics
physics is an independent acceptance dimension between an impressive video and a decision-useful world model. Fix actions and initial conditions, then measure response, drift, physical constraints, planning gain, and real-world transfer separately.
planning utility
planning utility is an independent acceptance dimension between an impressive video and a decision-useful world model. Fix actions and initial conditions, then measure response, drift, physical constraints, planning gain, and real-world transfer separately.
sim-to-real
sim-to-real is an independent acceptance dimension between an impressive video and a decision-useful world model. Fix actions and initial conditions, then measure response, drift, physical constraints, planning gain, and real-world transfer separately.
Build and verify
Build an evaluation matrix across Genie, JEPA, Marble, and Cosmos
- Predict: write the expected output, trend, or failure before running code.
- Build: implement only the minimum components needed to answer the question.
- Verify: compare with a baseline or trusted implementation; save seeds, parameters, and raw outputs.
- Transfer: change one shape, dataset, scale, or workload condition and explain whether the conclusion still holds.