Future Anchored Verification and Online Recovery for World Action Models
Organizations: National University of Singapore
Abstract
World action models (WAMs) have emerged as a promising paradigm for robotic manipulation. They act by first predicting how a task should be performed and then decoding the actions from that future. However, the remaining actions are invalid once execution drifts from the prediction. Simply replanning from the already out of distribution state rarely restores what the task still requires; existing execution monitors decide when to stop, but not what to restore. We observe that the answer is already in hand: the future the WAM predicted before acting depicts exactly the states it intended to pass through. We introduce FAVOR (Future Anchored Verification and Online Recovery), a lightweight framework that keeps these predicted frames as anchors and uses them for verification and recovery. An Anchor Verifier compares each observation with its anchor, together with the executed actions, to flag deviations that break the task. Anchor-Guided Recovery uses a vision-language model to turn the flagged anchor into a short corrective instruction. Under strengthened instruction guidance, the WAM executes this instruction to return to the intended future. It then resumes the task. FAVOR raises the task success of the base WAM from 97.85% to 98.10% on LIBERO and from 72.60% to 72.98% on LIBERO-Plus without modifying the policy.
Figures & tables
| LIBERO | LIBERO-Plus | |||||||
| Success (%, ) | Failure (%, ) | Success ( ) | ||||||
| Method | Spatial | Object | Goal | Long | Avg. | Rate | Rel. | Avg. |
| Fast-WAM-IDM | 98.20 | 98.90 | 98.30 | 96.00 | 97.85 | 2.15 | – | 72.60 |
| + Verifier, direct replanning | 98.30 | 99.20 | 98.20 | 96.00 | 97.92 | 2.08 | 3.3% | 72.60 |
| + Verifier, correction w/o anchor | 98.30 | 99.30 | 98.40 | 96.10 | 98.02 | 1.98 | 7.9% | 72.53 |
| FAVOR (Anchor-Guided Recovery) | 98.30 | 99.40 | 98.40 | 96.30 | 98.10 | 1.90 | 11.6% | 72.98 |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Meaning | Value |
|---|---|---|
| Predicted actions per chunk | 16 | |
| Executed actions per chunk before replanning | 16 | |
| Actions between consecutive anchors | 4 | |
| Alarm threshold of the fixed operating point | 0.5 (LIBERO), 0.9 (LIBERO-Plus) | |
| Maximum length of the correction window (actions) | 64 | |
| Guidance scale during recovery | 5 |