For control, the useful object is an action-conditioned, updateable representation that predicts task-relevant consequences of interventions and supports planning, revaluation, and transfer.

Beyond pixel prediction

A model can generate visually convincing futures while failing to preserve the structure that matters for action. An embodied agent needs to distinguish what it can control, what changes independently, and which relations persist across tasks.

One direction is to separate latent state into context that should remain stable within a trajectory and factors that change over time. The dynamic component should preserve temporal and controllable structure; the combined representation should retain enough information for reliable action without requiring pixel-perfect prediction.

A model for intervention

The test of a world model should therefore include counterfactual action and adaptation, not only rollout appearance. Can it predict how different interventions change contact, motion, and task outcomes? Can it update when the environment violates an assumed relation?

Open question. How can invariant context and controllable dynamics emerge through self-supervision, without goals that narrow generalization?

This is a working note. Discuss the idea or return to the writing archive.