Considering Context: When World Models Need Context Encoders
Organizations: King AI Labs, Microsoft Gaming
Abstract
Methods for generalization in model-based reinforcement learning typically assume that an agent cannot recover the latent context governing the environment dynamics from its own experience, and therefore supplies it externally. We formalize and test this assumption with \emph{predictive sufficiency}, which quantifies what access to the context adds to next-step prediction under the visitation distribution an agent induces, and separates that quantity into a history-recoverable part, a residual requiring the true context, and the deficit added by a finite model. We classify context-aware algorithms by the predictive risk their conditioning set can target and demonstrate across environments of increasing identification difficulty that the headroom does not follow the MDP class. The same task under different priors leaves predictive headroom in one setting and nothing distinguishable from zero in another, where the agent's behavior implicitly identifies the context and any benefit of such a mechanism cannot be attributed to missing information. Where headroom persists, the learned state exposes it only partially, and adding the true context still lowers the risk. Our contribution is a practical criterion for matching contextual mechanisms to the information available to them, estimated from the ordinary trained agent without a reference policy.
Figures & tables
| State | Setting | |||
|---|---|---|---|---|
| Discrete | Wind ( ) | 0.871 | 0.620 | |
| Discrete | Resistance ( ) | 0.062 | 0.059 | |
| Continuous | Attractor ( ) | 1.006 | 1.006 |
| Class | Conditioning set (bound) | Sufficient when | Methods |
|---|---|---|---|
| Memoryless | ( ) | decodable from | Base probes |
| Transition window | last transitions ( ) | recent evidence | DALI-S , CaDM, UP-OSI, GrBAL, PEARL, IIDA |
| Recurrent | within-episode summary ( ) | evidence within the episode | DreamerV3 , ReBAL, VariBAD |
| Cross-episode chain | cross-episode summary ( ) | episodes correlated | Carried state , LILAC , RL 2 , L2RL |
| Ground-truth context | realized ( ) | always | cRSSM-D , cRSSM-C |
| Setting | Environment | Context process | Difficulty level | Tests |
|---|---|---|---|---|
| Calibration | Point mass | persistent two-state chain, i.i.d. and short-episode variants | within/across episodes | P3 (control) |
| Identifiable | CARL Walker | i.i.d. per episode | current state and action | P1, P4 |
| Hard | CARL Walker | i.i.d. per episode | episode history | P2, P4 |
| Drift | RWRL Walker | non-i.i.d. Markov chain across episodes | across episodes | P3 |
| Identifiable | Hard | |||||
|---|---|---|---|---|---|---|
| Method | Return | Headroom | Resid. gain | Return | Headroom | Resid. gain |
| DreamerV3 | ||||||
| DALI-S | ||||||
| LILAC-style | ||||||
| cRSSM-D | ||||||
| Start transient | Start prior loss | |||||
|---|---|---|---|---|---|---|
| Method | Low | Mid | High | Low | Mid | High |
| DreamerV3 | ||||||
| Carried state | ||||||
| DALI-S | ||||||
| LILAC-style | ||||||
| cRSSM-D | ||||||
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Agent setting | Value |
|---|---|
| Backbone | DreamerV3, block-GRU |
| Parameters (Walker baseline) | M ( M transition) |
| Deterministic state | |
| Stochastic state | categorical |
| RSSM hidden width | |
| Encoder / decoder | layers, width |
| (a) Probe calibration | ||
|---|---|---|
| Quantity | Exact | Est. |
| Context score gap | ||
| History recovery | ||
| Shuffled null | ||