Organizations: Centre for Innovation and Precision Eye Health, Yong Loo Lin School of Medicine, National University of Singapore, Singapore 119228, Singapore · Department of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore 119228, Singapore · Bioengineering Program, Biomedical Sciences Division (BioMed), King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia · Institute of Advanced Intelligence and Computing (IAIC), Agency for Science, Technology and Research (A*STAR), Singapore 138632, Singapore · Singapore Eye Research Institute, Singapore National Eye Centre, Singapore 169856, Singapore · Ophthalmology & Visual Sciences Academic Clinical Program (EYE ACP), Duke-NUS Medical School, Singapore 169856, Singapore
World models enable agents to learn and plan in imagination, but predictions beyond their experience can become unreliable and mislead decisions. Existing uncertainty estimates derived from predictions can remain overconfident on unfamiliar state-action pairs. We propose the Lucid World Model (LucidWM), which learns doubt from experience and propagates trust through imagination. By integrating Subjective Logic into categorical latent transitions, LucidWM distinguishes predicted outcomes from their evidential support and assigns each transition a degree of doubt. The complement of this doubt defines transition-level trust, which accumulates multiplicatively along imagined trajectories to reweight returns for policy learning and guide action selection. Uncertainty estimation requires no additional parameters or forward passes. Evaluated on four base world models against seventeen uncertainty readouts, LucidWM detects environmental changes and signals uncertainty during action-corrupted rollouts. In a controlled navigation case study, acting on trust reduces the number of steps required to reach the goal from 362 to 190. Fifteen demonstration videos show how LucidWM doubts its dreams and acts on that doubt. Videos are available at https://lucidwm.github.io.
Figures & tables
Figure 1: Our world model with doubt. The doubt is read from the same head that generates its imagination, with no added parameter. The dark line takes one real fall as an example: the model reports its doubt eight steps before the fall and stays above its alarm line throughout it. We show that it reports all twenty consecutive real falls in the same way, every one before the fall begins. The usual readout, the entropy of the prediction, reports none of them.
Figure 2: Same prediction, different experience. Top: a scene and an action never experienced together are predicted as sharply as a pair that was (real frames: at a fall’s onset, and 30 steps before it). Bottom: ours reads the Dirichlet behind the input, whose strength and doubt separate the two; the entropy of the output cannot. Bars: each readout over its own alarm line.
Figure 3: Doubt is learned from experience and carried as trust. (a) A standard world model; (b) LucidWM, changes in grey and blue, doubt in orange. 1) The head’s evidence e yields a doubt u in place of the fixed ε ; the discipline Lu lets evidence recede toward the vacuous opinion where the dynamics loss does not hold it up. 2) Doubt sets each imagined step’s trust τk , compounded into Ck ; Rτ keeps λτk of each step where Rλ keeps λ . 3) A course whose trust collapses is vetoed.
(A) After observation
free readouts
fitted heads
multiple forwards
base
task
ours 1×
base 1×
entropy 1×
KL 1×
recon. 1×
1-step 1×
RND +.18
RND-e +.38
evid. +.12
latent +.62
self 3 fwd
deep 3×
DreamerV3
maze
0.77
0.24
0.31
0.24
0.06
0.10
0.39
0.27
0.48
0.33
0.21
0.21
R2-Dreamer
0.83
0.37
0.36
0.52
–
0.50
0.45
0.66
0.30
0.33
0.41
–
EMERALD
0.96
0.72
0.40
0.27
0.05
0.11
0.22
0.20
0.85
0.83
0.26
–
OC-STORM
0.82
0.75
0.75
0.37
0.09
0.16
0.53
0.46
0.57
0.59
0.24
–
(B) Before decision
free readouts
fitted heads
multiple forwards
Table 1: (A) After observation: AUROC of the readout, frames of the opened room against familiar ones; 0.5 is chance. (B) Before decision: lift ρ1/ρ0 , the readout under corrupted actions over the readout under true ones; 1 is no response. Under each column its cost relative to the base ( +c : heads with c times its parameters; m fwd: m forward passes). Means over five seeds; bold is the best in a row; grey numbers move the wrong way, falling as the model leaves its experience; a dash (–) is a readout the base cannot provide. Each test uses eleven of the seventeen readouts; reconstruction and one-step error need the arriving frame, so appear in (A) only.
Figure 4: Ours doubts the departure; base does not. Top: experienced states and 48 random-action rollouts coloured by doubt. Bottom: doubt along the rollouts.
Figure 5: The alarm fires where the world changed. (a) A drive into the repainted room. (b) A patrol that twice faces the newly opened room. Doubt in each model’s own units; dashed line: alarm line; sand: frames in the changed region, where the gold-outlined pictures were taken.
Figure 6: Learning more, doubting less. Rows: checkpoints of one run; columns: imagined steps. Tiles: pixel error, dark right, bright wrong. Shaded: chaotic horizon, never learned.
Figure 7: Trust tells which imagined steps to keep. Error removed by dropping the worst fifth of steps: (a) three DMC tasks, (b) cheetah from pixels. Pale: over all depths; solid: within each depth; whiskers: bootstrap 95% intervals.
Figure 8: One veto, half the steps. Top: steps to the armour. Middle: both routes from the repainted dead end (sand) to the armour (star), and the two views at step 183, seven steps before ours arrives. Bottom: doubt along ours and the entropy of base, each on its own p99 alarm line; circles: where ours turns and base leaves.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
symbol
meaning
defined
standard
World model and imagination
xt,st,at
observation, latent state, action at time t
Sec. 2
G,K
number of categorical variables of the state; classes per variable
Sec. 2
zt+1
one categorical variable of the next state
Eq. ( 1 )
ℓ(st,at)
the transition head’s K logits (per variable)
Eq. ( 1 )
p(zt+1∣st,at)
the head’s prediction
Eqs. ( 1 ), ( 4 )
Appendix
Table 2: Symbols of Secs. 2 – 4 . Per variable: defined for each of the G categorical variables, whose index g is dropped when clear. The last column gives the value that recovers the standard world model.
stage
standard
LucidWM
fallback
prediction, Eq. ( 4 )
(1−ε)softmax(ℓ)+ε/K
(1−u)p^+u/K , u from S
uniform 1/K
training, Eq. ( 5 )
fit only
fit and evidence discipline
the vacuous opinion
posterior, Eq. ( 6 )
a second network q
e⊕eobs
none: doubt never rises
return, Eq. ( 7 )
share λ per step
share λτt+1 , τ from uˉ
critic v
decision, Eq. ( 10 )
reward counted in full
reward weighted by Ck
zero
Appendix
Table 3: What LucidWM changes in a categorical world model. Every readout is a function of the head’s logits; the standard model is recovered by u≡ε and τ≡1 .
step
result
where
measure
the doubt u=W/(W+S) sees the evidence only through its total S
Eq. ( 3 )
ideal
in a table updated by Bayes’ rule, S is the joint count; an untaken pair has u=1
Prop. 2
obstacle
neither a readout of the prediction nor the fit can see S
Prop. 3
remedy
the discipline keeps the most doubtful opinion with its mean; u=1 if untaken
Prop. 4
payoff
the doubt sets the expected error of the prediction
Prop. 5
Appendix
Table 4: The argument of App. C , one result per step.
Figure 9: One prediction, many evidence totals , for the example pair. (a) Dirichlets with the same mean P (dot) and totals S=8 , 18 and 198 . (b) Change from S=8 along this family: the fit and every readout stay constant (Prop. 3 ), while the discipline grows with the total (Lemma 1 ); among opinions with this mean it prefers the smallest total, S=8 , the count. (c) The expected error falls with the total (Prop. 5 ); at S=198 , the total a standard head asserts, it is eighteen times smaller than at the count.
step
result
where
start
at the real state, fusing in an observation never raises the doubt
Prop. 6
alternative
carried in the state, the doubt of a step is lost one step later
Prop. 7
choice
carried in the return, trust compounds, as SL trust discounting does
Prop. 8
safety
the lucid return keeps the critic’s fixed point
Thm. 1
payoff
errors that follow a doubtful step are discounted, where it pays
Prop. 9
action
doubt ranks candidate actions before any is taken
Prop. 10
Appendix
Table 5: The argument of App. D , step by step.
Figure 10: Trust in the state loses the lineage; trust in the return compounds , for the example rollout. (a) Trust at each imagined step n : deduced through the state, (1−u~n)/(1−ε) with σ=0.009 (Prop. 7 ), and carried in the return, Cn (Prop. 8 ). (b) Weights of the one- to eight-step returns under the λ -return and the lucid return (Lemma 2 ).
base
dynamics
G×K
W
H
DreamerV3
recurrent, pixel decoder
32×32
2
15
R2-Dreamer
recurrent, no reconstruction
32×16
1
15
EMERALD
masked transformer, spatial
4×4 cells of 32×32
2
15
OC-STORM
transformer
32×32
2
16
Appendix
Table 6: The four bases; ε=0.01 and λ=0.95 on all. G×K : categorical variables and classes of the latent state; W : prior weight; H : imagination horizon.
experiment
bases
environments
result
main text
after observation, Tab. 1 (A)
all four
maze
first on every base
before decision, Tab. 1 (B)
all four
DMC, Crafter
first on all 16 rows
leaving experience, Fig. 4
DreamerV3
walker
9 in 10 over the line; base flat
two changed rooms, Fig. 5
DreamerV3
maze
alarm in both; base in neither
learning, Fig. 6
DreamerV3
cheetah
doubt falls with the error
Appendix
Table 7: Every experiment in the paper and what it finds. DMC in Table 1 (B): walker, cheetah, pendulum.
Figure 11: Into the new room, and into it again. (a) One walk: views outside, inside (gold), back and inside again, and the doubt of Ours and Base in alarm units; sand: in the room; red: alarms. (b) Walks on which each readout raises its alarm inside the room.
Figure 12: One pass against three models. (a) AUROC of the new room against relative parameters, as in Table 1 (A); grey band: below chance, where a readout reads the new room as more familiar than the maze. (b) Walks through the opened door on which each readout raises its alarm, with its cost; top: walk 10 at −14 , 0 and +25 steps from entry, gold where Ours raises its alarm, joined to its dot.
Figure 13: Twenty falls, twenty alarms before the fall. (a) Doubt (blue) and the entropy of the same prediction (grey) around the onset of each of twenty consecutive falls, in alarm units; dashed: alarm line; red: onset of the doubt’s alarm that holds into the fall; sand: the fall. (b) Falls whose alarm has fired by each step. Top: fall 5, the fall of Fig. 1 , at −20 , −6 , 0 and +5 steps; gold: step −6 , where its alarm holds; the doubt first crosses the line at −8 , which Fig. 1 counts as its lead.
Figure 14: Nightfall. Views from fifteen steps before nightfall to thirty-five after (gold: at the alarm), above the doubt of Ours and Base in alarm units; sand: the night; strip: the fading light.
Figure 15: Four bases, four alarms. Doubt of Ours and Base on each base as the agent approaches the new room, in alarm units; sand: the room in view; red: the first alarm of Ours. Top: the route plotted in (b) and (d) at −14 , 0 and +11 steps from entry, gold at the alarm of (b), joined to its dot; (a) and (c) pass the same door.
Figure 16: The worse the actions, the higher the doubt. Lift against the share of actions corrupted; thin line at one: no response. (a, b) OC-STORM, median over 300 starts (band: interquartile range of Ours); insets: imagination under true and under corrupted actions. (c, d) EMERALD against MC dropout, Laplace and a snapshot ensemble, on a log scale; insets: a real frame of each task.
Figure 17: Any alarm line, any checkpoint. (a) Rules, placement of the line (rows) by frames required above it (columns), under which each readout raises its alarm in the repainted room. (b) AUROC of the opened room at each of 38 checkpoints of training; dashed: chance. Top: each change seen from one viewpoint, as trained and as changed.
Figure 18: Six checkpoints, one future. As Fig. 6 , at all six checkpoints of the run. Right: error, doubt and the entropy of the same prediction, each averaged over the 24 steps, relative to the first checkpoint.
Figure 19: Trust fades faster where the dream drifts. Real (top) and imagined (bottom) futures under one action sequence. (a) Walker; gold: the imagined walker has fallen; the pixel error between the two futures. (b) Crafter; gold: lava in the real world; trust of this rollout against the median over thirty calm starts.
Figure 20: Trust across ten runs, and two hyperparameters. (a) Error removed by dropping the worst fifth of imagined steps, ranked within each depth as in the solid bars of Fig. 7 , averaged over ten runs across walker, cheetah, finger and cartpole; top: the four tasks. (b) The same, ranked over all depths on cheetah from pixels as in Fig. 7 b, with trust compounded over the last w steps; w=15 : the whole rollout, as in LucidWM. (c) Return on walker against the discipline weight βu . Whiskers: bootstrap 95% intervals.
Figure 21: Fewer deaths, faster learning. (a) One agent choosing among imagined futures scored with and without trust; median life: per seed, averaged over 50 seeds. (b) Evaluation return over training; dots and lines: where each first reaches 8. (c) The official Crafter score of the 22 achievements over the last 100 evaluation episodes; whiskers: bootstrap 95% intervals over episodes; tiles: from the scored episodes, Ours placing a furnace and Base a table.
Figure 22: Vetoes fire where the world changed. Where every veto fired on (a) the opened map and (b) the repainted dead end; sand: the changed rooms; rings: vetoes outside them. Insets: the agent’s view at a veto, gold and joined to its dot; in (b), also the repainted room at step 8. The veto at step 24 in (b) is that of Fig. 8 .