🌍 World Models
Learn to predict action-dependent futures, plan with a small dynamics model, and test whether imagined rollouts remain useful in the real environment.
On this page
Before you start
Reinforcement Learning Basics introduces states, actions, rewards and policies. Deep Learning Basics explains learned predictors, and Probability & Statistics helps with uncertain futures.
You will learn to separate dynamics, observations and action selection, perform a short planning rollout, and explain why visual realism alone does not establish a useful simulator.
The problem: try an action before taking it
Imagine a robot sliding an object toward a target. A policy can map its camera image directly to a push. A world model gives it another option: predict what several possible pushes would do, compare the outcomes and execute the first action of a promising plan.
Broadly, a world model represents or predicts aspects of an environment. For control, an action-conditioned dynamics model predicts how the environment changes under an action. Some world models predict pixels, others object states or learned features. They need not represent a 3D scene, obey physics exactly, or generate a human-viewable video to be useful.
The distinction from an LLM is the modeled variables and purpose, not whether the architecture is a transformer or whether the inputs are tokenized. A text model is not automatically a reliable action-conditioned simulator; a video tokenizer does not prevent a model from learning dynamics.
Name the pieces
| Symbol | Meaning | Example |
|---|---|---|
| s_t | Environment state at time t | Object position, velocity, contact state |
| o_t | Observation | Camera image or sensor reading |
| a_t | Action applied during the next interval | Motor command or acceleration |
| z_t | Learned representation of available history | Latent features from observations/actions |
| f | Dynamics: state/action to next-state distribution | Predict where the object moves |
| g | Observation model: state to observation | Render a camera view |
| π | Policy: available information to action | Choose a push |
In a partially observable setting, one image may not reveal velocity or a hidden object. A history encoder or belief state summarizes what is known; a latent vector is not automatically a calibrated belief distribution.
Dynamics prediction, rendering and planning are different functions. A planner searches or optimizes actions using a model and an objective. It is not generally the inverse of a renderer: many states can produce similar images, and many action sequences can reach a goal. A VLA can act directly without running an explicit search.
A two-second worked plan
Use an idealized one-dimensional object with position x in meters, velocity v in meters/second, action a in meters/second², and time step Δt in seconds. Constant-acceleration dynamics are:
x_next = x + vΔt + 0.5aΔt²
v_next = v + aΔt
These equations are a known reference model for the exercise; a learned world model would approximate the transition from data. Start at x=0 m, v=1 m/s, with Δt=1 s. The goal is to stop at x=1.5 m after two steps.
| Actions a_0, a_1 (m/s²) | State after step 1 | State after step 2 |
|---|---|---|
| 0, 0 | x=1 m, v=1 m/s | x=2 m, v=1 m/s |
| −1, 0 | x=0.5 m, v=0 m/s | x=0.5 m, v=0 m/s |
| 0, −1 | x=1 m, v=1 m/s | x=1.5 m, v=0 m/s |
Define a dimensionless terminal cost:
J = ((x_final − 1.5 m)/(1 m))² + 0.1 × (v_final/(1 m/s))²
The three costs are 0.35, 1 and 0. The third plan wins among these candidates. Normalizing the terms makes their units explicit; adding raw squared meters to squared meters/second would hide a units mismatch.
In model-predictive control, execute only the first action, observe the new state, then plan again. If the object actually reaches x=0.9 m with v=0.8 m/s, use that observation. Do not blindly follow the rest of a rollout that assumed x=1 and v=1. Prediction error is why feedback matters.
How a learned model gets trained
Collect trajectories containing observations, actions, timestamps and relevant outcomes. Align action intervals with sensor observations; a one-frame timing error can teach the wrong transition. An encoder produces z_t, a transition model predicts z_next or a distribution over futures, and optional heads predict rewards, continuation or reconstructed observations.
Losses depend on the representation: state regression with meaningful units, observation likelihood/reconstruction, or prediction of target embeddings. Stochastic dynamics matter when information is hidden or several futures are plausible. A deterministic average of “pass left” and “pass right” may predict an impossible collision through the middle.
Split evaluation by trajectories, environments or collection sessions, not adjacent video frames that nearly duplicate training examples. Logged actions may cover only an expert's narrow behavior. Predicting those logs accurately does not establish counterfactual accuracy for unfamiliar actions; investigate support and confounding in data collection.
Pixels, latents and geometry serve different jobs
| Representation | Useful for | What it does not guarantee |
|---|---|---|
| Video/pixels | Visual observation generation and human inspection | Stable geometry, valid contacts or action causality |
| Latent features | Compact prediction and planning | Retention of every task-relevant detail |
| Meshes and physical state | Contact queries and explicit dynamics | Photorealism or correct physical parameters |
| Gaussian splats / radiance fields | View synthesis from a persistent scene representation | Collision surfaces, mass, friction or articulated dynamics |
A splat is an appearance primitive. A collision mesh is a separate geometric artifact, which may itself need simplification or convex decomposition for a physics engine. Neither tells you the correct mass, friction or compliance. Video models may maintain latent state or memory; streamed pixels do not prove that they have none.
Latent approaches such as JEPA predict learned representations rather than reconstruct every pixel. That can avoid modeling some irrelevant detail, but task-relevant geometry can also be lost. Examine the training objective and collapse-prevention recipe; an encoder that outputs a constant would make some naive prediction objectives trivially easy. Not all video representation models are action-conditioned without additional training.
Planning can exploit model errors
Repeated prediction feeds earlier errors into later predictions. Even a small one-step error can move a trajectory into states absent from training. A planner may actively seek a model's optimistic mistakes because those look like high reward.
Use short receding horizons, model ensembles or other uncertainty estimates, constraints, real-data correction and evaluation outside the training distribution. Uncertainty estimates can also be wrong under shift. More samples from the same flawed simulator do not establish that its physics is correct.
World models can supply imagined trajectories for policy learning, as in Dreamer, or evaluate candidate action sequences at inference. Model-free imitation can train a VLA directly from demonstrations, so a world model is optional. Synthetic observations and rewards need validation; their labels are simulator outputs, not automatic ground truth about reality.
Evaluate the job the model is meant to do
Measure action-conditioned one-step and multi-step prediction, object persistence, geometric/contact consistency, intervention response and downstream policy success. Record position error in meters, velocity error in meters/second, collision rates with a defined denominator, and latency relative to the control interval.
Visual metrics such as PSNR, SSIM or FVD measure selected reconstruction/distribution properties. They can favor realistic-looking clips that respond incorrectly to an action. Conversely, a compact latent model can help control without producing attractive video. Evaluate on controlled real tasks or a trusted reference where feasible, and report simulation validity failures separately from policy failures.
Check yourself
Use x=0 m, v=1 m/s, a=−1 m/s² and Δt=0.5 s. What next state should the reference dynamics produce?
Solution: x_next = 0 + 1×0.5 + 0.5×(−1)×0.5² = 0.375 m; v_next = 1−0.5 = 0.5 m/s. Using the one-second displacement would be a timestep error, not a small model-quality difference.
Where to go next
Vision-Language-Action Models connects observations and instructions to robot commands. Evaluation & Benchmarking helps design comparisons whose scores match the intended task.
References
- World Models: learned representations, dynamics and controllers.
- DreamerV3: learning behavior using imagined trajectories.
- V-JEPA 2: latent video prediction and action-conditioned planning.
- 3D Gaussian Splatting: an appearance representation for view synthesis, distinct from a physics simulator.