World-action models use predicted visual futures to condition robot actions, yet execution feedback can invalidate parts of a prediction while leaving its task structure useful. We propose Revisable Temporal Planning (RTP), which maintains the visual future as a persistent action condition and revises it after feedback. Its central mechanism is a learned revision bridge: it resumes an intermediate state saved during visual generation and adapts its continuation to current observations. Visual and action supervision connect this revision to subsequent control. Time-aware history supplies observed evidence, and an adaptive policy selects retention, bridge revision, or fresh replanning from new noise before decoding the next action. On RoboMME and RMBench, RTP achieves task-averaged success rates of 48.6% and 84.8%, respectively. Matched comparisons support learned continuation; estimated checkpoint-source and action-prefix effects are positive but less precisely resolved. These results connect feedback-driven visual-plan revision to closed-loop task performance. Project Page: https://PLACEHOLDER.github.io/RTP/
Figures & tables
Figure 2 : The RTP feedback loop. The upper row illustrates a fresh plan root. At each feedback boundary, update factual history, compare prediction with feedback, and run one selected visual update. Save the accepted record and decode a new action from its next prefix and current facts.
Figure 3 : Prediction and execution. Base-model continuation under a joint-target hold at samples 52–59. Prediction (a) and execution observation (b) disagree at 60; the updated prediction for 76 (c) better matches the approach progress in the subsequent observation (d).
Method
Success (%) ↑
Vision–language–action models
π0.5 ( Physical Intelligence et al., 2025 ; Dai et al., 2026 )
17.9
MME-VLA (TTT–Modul) ( Dai et al., 2026 )
22.0
MemER ( Sridhar et al., 2025 ; Dai et al., 2026 )
42.4
MME-VLA (FrameSamp–Modul) ( Dai et al., 2026 )
44.5
World–action models
Table 1: Task-averaged success on RoboMME.
Method
Success (%) ↑
Vision–language–action models
π0.5 ( Physical Intelligence et al., 2025 ; Chen et al., 2026b )
10.4
X-VLA ( Zheng et al., 2025 ; Chen et al., 2026b )
9.8
Mem-0 ( Chen et al., 2026b )
42.0
World–action models
Fast-WAM ( Yuan et al., 2026 ; Yang et al., 2026 )
5.9
Table 2: Task-averaged success on RMBench.
Table 5
Action facts
Action prefix
Success (%)
Prefix effect (pp)
95% CI
Previous
Retained
41.3
1.8
[-0.9, 4.4]
Previous
Revised
43.0
1.8
[-0.9, 4.4]
Current
Retained
44.9
2.5
[0.0, 5.0]
Current
Revised
47.4
2.5
[0.0, 5.0]
Table 5: Action-input ablation under the Fixed bridge-10 policy. Prefix effects compare revised minus retained using unrounded counts; intervals are paired.
Figure 4 : Paired success differences on RoboMME. Dots and bars show differences and paired 95% intervals for the labeled policy contrasts, which are not additive. Sources: Tables 10 , 6 , and 5 .
Condition
I (%)
O (%)
N (%)
N–O (pp)
95% CI
Nominal
43.6
44.6
47.4
2.8
[-0.1, 5.5]
Four-sample hold
43.6
44.6
47.4
2.8
[-0.1, 5.5]
Eight-sample hold
43.0
44.2
47.1
2.9
[0.0, 5.6]
Table 6: Fixed-depth source comparisons under actuator holds, 800 keys/condition. I/O/N use independent-noise reconstruction, root-noise reconstruction, and saved checkpoints, respectively, with ten-interval continuation and fresh at exhaustion; difference and paired interval: N minus O.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Quantity
RoboMME
RMBench
Future / consumed groups
4/1
4/1
Native samples per block J
4
16
History budget BH
60
72
Reset anchor quota
1
1
Recent nonanchor quota
12
16
Task-reference budget BM
0 or 32
0
Appendix
Table 7: History and model configuration. History-group budgets include all views; task references have a separate budget.
Task
Fresh
RTP
Rescues
Regressions
Δ [95% CI]
T01
13
13
–
–
0.00 [–]
T02
12
16
7
3
8.00 [-4.0, 20.0]
T03
11
14
4
1
6.00 [-2.0, 14.0]
T04
19
24
9
4
10.00 [-4.0, 24.0]
T05
16
18
5
3
4.00 [-6.0, 16.0]
T06
23
33
13
3
20.00 [6.0, 34.0]
Appendix
Table 8: Per-task RoboMME comparisons, 50 paired keys/task. Successes, rescues, and regressions are counts; differences and intervals are in percentage points. – denotes unreported paired statistics.
Task
Fresh
RTP
Rescues
Regressions
Δ [95% CI]
R01
61
63
–
–
2.00 [–]
R02
79
83
9
5
4.00 [-3.0, 11.0]
R03
77
81
6
2
4.00 [-1.0, 10.0]
R04
73
82
11
2
9.00 [2.0, 16.0]
R05
79
88
10
1
9.00 [3.0, 15.0]
R06
80
86
8
2
6.00 [0.0, 12.0]
Appendix
Table 9: Per-task RMBench comparisons, 100 paired keys/task. Rescues and regressions are counts; differences and intervals are percentage points. – denotes unreported paired statistics.
Contrast
Difference
95% interval
Foundation configuration
6.88
[4.13, 9.63]
Extra feedback fitting
3.12
[0.50, 5.63]
RTP vs. matched Fresh-20+C
5.125
–
Learned vs. zero, Fixed bridge-10
5.75
[3.00, 8.50]
RTP vs. Fixed bridge-10
1.250
–
RTP vs. binary retain/fresh
2.250
–
Appendix
Table 10: Paired policy and intervention contrasts, in percentage points. Overlapping contrasts are not additive module contributions. – denotes unreported intervals.
Policy
Visual steps
Success (%)
Call (s)
RTP diff. (pp)
Fresh-5
5.00
34.25
0.976
14.38
Fresh-5+C
5.00
36.50
0.991
12.13
Fresh-10
10.00
38.12
1.066
10.50
Fresh-10+C
10.00
41.75
1.081
6.88
Fresh-12+C
12.00
42.50
1.123
6.13
Fresh-20+C
20.00
43.50
1.275
5.13
Appendix
Table 11: Fresh integration schedules and RTP. Visual work includes structural refreshes; differences are RTP minus row.
Policy
Success (%)
Mean steps
Call (s)
Fixed bridge-5
45.62
8.49
1.057
Fixed bridge-10
47.38
12.33
1.115
Binary retain/fresh
46.38
11.10
1.113
RTP
48.63
7.78
1.051
Appendix
Table 12: Fixed-depth, binary retain/fresh, and adaptive updates. Each policy includes exhaustion refreshes.
Fitting
Nominal
Hold 4
Hold 8
Default RTP
69/4822 (1.43%)
91/4542 (2.00%)
94/4252 (2.21%)
Without age/length
181/4726 (3.83%)
176/4392 (4.01%)
219/4141 (5.29%)
Matched parent refit
74/4838 (1.53%)
91/4533 (2.01%)
93/4257 (2.18%)
Recurrent fitting
24/4944 (0.49%)
25/4666 (0.54%)
24/4376 (0.55%)
Appendix
Table 13: Recurrent discrepancy checks on 320 episodes/condition, with common frozen tolerances. Entries are exceedances / selected reuse (percent).
Policy
Success (%)
Mean steps
Difference [95% CI]
Default RTP
48.63
7.78
0.00 [0.00, 0.00]
Without age/length
48.00
7.80
-0.625 [–]
Matched parent refit
49.38
7.79
0.750 [–]
Recurrent fitting
49.50
7.75
0.875 [–]
Recurrent − matched refit
–
–
0.125 [-2.50, 2.75]
Appendix
Table 14: Complete-policy recurrent-fitting controls on 800 reset keys. The first four differences are row minus default RTP; the final row is their direct recurrent-minus-matched contrast, not an additional policy. Differences and intervals are in percentage points; – denotes unreported intervals.
Boundary
Reached
Reused
Fresh
Exceedances
Exceedance (%)
1
2444
2168
276
27
1.25%
2
2070
1593
477
23
1.44%
3
1516
1061
455
19
1.79%
4
1016
0
1016
0
N/A
Total
7046
4822
2224
69
1.43%
Appendix
Table 15: Diagnostics by feedback boundary within each root. Exceedances count selected reuse with either distance above tolerance. N/A denotes the undefined zero-reuse ratio 0/0 .
Candidate
Contact MAE
Order (%)
Visual error
Retain
3.02
59.69
0.1251
Fresh
2.01
74.38
0.0893
Reconstruction-10
2.43
70.31
0.0955
Bridge-10
1.63
79.06
0.0691
Appendix
Table 16: Prediction against each candidate’s own action-conditioned continuation, 320 complete windows per branch; contact mean absolute error (MAE) is in native samples.
Condition
Zero (%)
Learned (%)
Difference (pp)
Nominal
41.62
47.38
5.75
Hold 4
41.62
47.38
5.75
Hold 8
40.75
47.12
6.38
Appendix
Table 17: Zero versus learned residual at fixed Fixed bridge-10. Source construction, depth, and action interface are matched; differences are learned minus zero.
Training loss
Success (%)
Visual error
Full–row (pp)
All terms
47.38
0.0840
0.00
No observed-visual loss
44.38
0.1206
3.00
No action loss
44.88
0.0805
2.50
No fresh regularizer
46.12
0.0961
1.25
Appendix
Table 18: Training-loss ablations at fixed Fixed bridge-10. Success and visual error have different targets; differences are full minus row.
Hold
Fresh
Fresh-20+C
Adaptive recon.
RTP
Difference (pp)
None
40.38
43.50
45.62
48.63
3.00
4
40.38
43.50
45.62
48.63
3.00
8
39.62
43.12
44.75
48.25
3.50
Appendix
Table 19: Complete-policy comparisons under actuator holds, 800 keys/condition. Differences are RTP minus Adaptive recon. Recon. denotes reconstruction from independent noise.
Condition / policy
Mean steps
Mean call (s)
4 / RTP
8.28
1.060
4 / Adaptive recon.
8.24
1.065
8 / RTP
8.88
1.071
8 / Adaptive recon.
8.95
1.078
4 / Fixed bridge-10
12.33
1.115
4 / Fixed recon.-10
12.32
1.117
Appendix
Table 20: Controller work under actuator holds. Noninitial calls use each policy’s own denominator. Recon. denotes reconstruction from independent noise.
Condition
Adaptive root recon. (%)
RTP (%)
Root recon. steps
RTP steps
Difference (pp)
Nominal
46.12
48.63
8.07
7.78
2.50
Hold 4
46.12
48.63
8.65
8.28
2.50
Hold 8
45.50
48.25
9.27
8.88
2.75
Appendix
Table 21: Original-root reconstruction and RTP. Differences are RTP minus Adaptive root recon.; step counts include all refreshes. Root recon. denotes reconstruction with the original root noise.
Delay D
Fresh-20+C
Fresh-10+C
Adaptive recon.
RTP
Difference (pp)
0
43.50
41.75
45.62
48.63
3.00
1
40.88
40.25
42.38
46.62
4.25
2
37.50
37.62
38.62
43.38
4.75
Appendix
Table 22: Activation-delay comparisons on 800 reset keys per delay. Differences are RTP minus Adaptive recon. Recon. denotes reconstruction from independent noise.
Policy
Success (%)
Mean steps
Matched diff. (pp)
Fresh-20+C
38.75
20.00
6.25
Adaptive recon.
41.25
8.94
3.75
RTP
45.00
9.03
0.00
Fixed recon.-10
39.50
12.31
3.75
Fixed bridge-10
43.25
12.32
0.00
Appendix
Table 23: Spatial-shift comparisons on 400 reset keys. Complete-policy differences use RTP; fixed-depth differences use Fixed bridge-10. Recon. denotes reconstruction from independent noise.
Setting
Fresh
Fresh-20+C
Adaptive recon.
RTP
RTP–Adaptive recon. (pp)
H=6,J=4
42.12
45.00
47.25
50.62
3.38
H=4,J=8
37.62
40.12
41.38
44.12
2.75
Appendix
Table 24: Horizon and sample-spacing comparisons on 800 reset keys per interface. Entries are success percentages. Recon. denotes reconstruction from independent noise.
Policy
Mean (s)
P50 (s)
P95 (s)
Fresh
1.248
1.248
1.304
Fixed bridge-5
1.057
1.007
1.281
Fixed bridge-10
1.115
1.088
1.268
RTP
1.051
1.010
1.294
Shared initialization
1.277
1.278
1.335
Appendix
Table 25: Controller-call durations. Policy rows contain noninitial calls; the last row contains the 800 shared fresh initializations. Quantiles use nearest rank.
Policy
Updates
Retain
Bridge-5
Bridge-10
Fresh
s/episode
Fresh
17403
0
0
0
17403
28.430
Fixed bridge-5
17882
0
13718
0
4164
24.902
Fixed bridge-10
17947
0
0
13769
4178
26.293
RTP
17922
6574
4407
2139
4802
24.817
Appendix
Table 26: Mode counts and episode cost on 800 episodes. Episode time includes initialization and all noninitial calls.