12.5 World models and video generation
You can follow the debate about what these systems learn.
Before:08. Large Language Models
World models and video generation aim at temporal and physical coherence — learned simulators with links to robotics and planning — and whether these systems learn physics or plausible-looking correlations is a live debate worth following honestly. It sits in the frontier module as the speculative edge. The judging error is the demo reel: selected clips prove selection, and evaluation here is genuinely unsolved, which is part of what makes the area worth watching.
Work through these
Video diffusion and temporal consistency
Generating video requires the frames to remain consistent with each other, which is a harder problem than generating any single one. Temporal consistency is the field's main technical difficulty.
Learned simulators and planning
Systems trained to predict what happens next can be used to plan, which goes considerably further than generating plausible video. Whether they genuinely do is the open debate.
Robotics and embodied AI links
The link to robotics is direct: a system that predicts consequences can choose actions. This is where generation stops being entertainment and becomes control.
Evaluation difficulties
Judging these systems is unusually hard because plausibility and correctness diverge, and human evaluation is expensive. Weak evaluation is why announcements of progress here are difficult to assess.
Sign in to keep your progress.
Free resources
We haven't checked most of these for screen reader use yet.
Links last checked 29 Aug 2026.
Stuck here?
Ask a mentor. A real person answers, and they can see exactly which topic you're on. Usually within a couple of working days.
Checking your session…
Topics shown in module order.