AquaMend: Minimal Re-probing and Conditional Rollback for Latent-Belief Failures in Embodied Agents
Organizations: University College London · Peking University · School of Electrical and Electronic Engineering, Nanyang Technological University · University of the Chinese Academy of Sciences · Beijing University of Posts and Telecommunications · University of Leeds · Fullive-AI
Abstract
Physical changes or sensing errors can invalidate embodied agents' task-relevant beliefs. AquaMend compares re-probing, rollback, and supported continuation on a probe-belief-action graph under an expected-loss objective covering sensing, physical recovery, and uncorrected failures. A joint posterior guides a one-step policy with conditional detection-power screening. The per-belief three-way optimum requires independence, separability, and fully resolving probes; the general policy has no global optimality guarantee. Across 32 paired scenarios in a self-constructed simulation benchmark, AquaMend recovers in 28/32 cases and reduces mean complete loss by 21.6% versus restart. Its paired loss difference from decision-theoretic troubleshooting (DTT) is not statistically significant after Holm correction. Against the all-candidate ablation, online decision time decreases by 12.3% overall but increases by 3.4% in the uncovered late stage.
Figures & tables
| Type | Symbol | Semantics | Cost |
|---|---|---|---|
| Probe | one latent-property measurement | redo | |
| Belief | latent proposition (e.g., “cup 3 is least full”) | residual | |
| Action | e.g., “pour into cup 3” | rollback |
| Mode | Definition | Complexity |
|---|---|---|
| Indep | for disjoint charged recovery operations | |
| Beam- | approximate using up to weighted joint failure hypotheses | |
| MC | average over posterior samples of the failure set |
| Family | Operation | Failed beliefs | Attribution |
|---|---|---|---|
| Swap cups | swap two containers’ poses/IDs | water level / target of swapped instances | External |
| Add water | add across a decision boundary | filled cup’s level and ranking | External |
| Sensor drift | bias / timestamp / mis-binding (world unchanged) | beliefs on the corrupted evidence | Self_Error |
| False alarm | brief occlusion / noise spike (no task-relevant failure) | False_Alarm |
| Granularity | Edge semantics | Expectation |
|---|---|---|
| Linear chain | stage dependency; rollback = suffix | benefit mainly from zero rollback on false alarms |
| Fine causal | explicit causal/evidence edges; rollback = minimal upper closure | rollback set itself is smaller |
| Method | (Holm) | ||
|---|---|---|---|
| AquaMend (ours) | 0.976 | 12.47 | – |
| Lookahead (depth 2) | 0.976 | 12.47 | 0.317 |
| Decoupled | 0.575 | 47.83 | |
| Greedy re-probe | 0.575 | 46.14 | |
| Restart | 0.333 | 53.16 |
| Method | [95% CI] | (Holm) | ||
|---|---|---|---|---|
| AquaMend (ours) | 28/32 | 71.356 | – | – |
| DTT | ||||
| No extra re-probe | ||||
| Full restart | ||||
| Linear chain |
| Method | Realized cost | Expected policy risk |
|---|---|---|
| Exact Adaptive | 8.560 | 9.253 |
| Lookahead (depth 2) | 8.565 | 9.255 |
| AquaMend (ours) | 8.594 | 9.290 |
| Decoupled | 11.008 | 11.756 |
| Greedy re-probe | 12.097 | 12.491 |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Mode | Specification | Control parameter |
|---|---|---|
| Indep | Uses posterior marginal belief-failure probabilities. For pairwise-disjoint rollback sets, . | Marginal probabilities |
| Beam- | Retains up to highest-probability joint belief-failure hypotheses, with candidate expansion along correlation edges, and approximates rollback risk using the retained hypotheses. | Beam width |
| MC | Estimates rollback risk from samples of the joint belief-failure configuration drawn from the posterior. | Sample count |
| Beliefs | Worst-case depth | Policy | Realized cost | Expected risk | Time (ms) | Memory (KiB) | |
|---|---|---|---|---|---|---|---|
| 3 | 2 | AquaMend Myopic | 0.800 | 3.396 | 3.396 | 0.512 | 22.4 |
| 3 | 2 | Lookahead (depth 2) | 1.000 | 3.248 | 3.248 | 0.705 | 22.6 |
| 3 | 2 | Lookahead (depth 3) | 1.000 | 3.248 | 3.248 | 0.744 | 39.3 |
| 3 | 2 | Exact Adaptive | 1.000 | 3.248 | 3.248 | 0.533 | 19.0 |
| 4 | 3 | AquaMend Myopic | 0.000 | 4.000 | 4.000 | 1.013 | 28.2 |
| 4 | 3 | Lookahead (depth 2) | 0.600 | 3.679 | 3.679 | 2.931 | 54.9 |
| Method | (Holm) | ||||||
|---|---|---|---|---|---|---|---|
| Oracle (reference) | 1.000 | 3.31 [3.21, 3.40] | 0.768 | 0.00 | 1.22 | 1.00 | |
| Restart | 0.799 | 19.52 [19.23, 19.83] | 3.00 | 3.00 | 0.00 | ||
| Blind retry | 0.252 | 23.70 [23.26, 24.10] | 2.00 | 0.00 | 1.00 | ||
| VeriTrace (rollback only) | 0.204 | 39.64 [39.05, 40.21] | 0.00 | 2.50 | 0.00 | ||
| Troubleshooting | 0.555 | 20.26 [19.69, 20.76] | 1.00 | 1.50 | 1.00 | ||
| Belief-space replan | 0.204 | 33.60 [32.99, 34.18] | 0.00 | 0.26 | 1.00 |
| Method | (Holm) | |||||
|---|---|---|---|---|---|---|
| Oracle (reference) | 1.000 | 7.83 | 0.00 | 2.98 | 0.000 | – |
| AquaMend (ours) | 0.976 | 12.47 | 2.88 | 2.97 | 0.001 | – |
| Lookahead | 0.976 | 12.47 | 2.88 | 2.97 | 0.001 | 0.317 |
| Decoupled | 0.575 | 47.83 | 3.61 | 1.83 | 0.403 | |
| Greedy re-probe | 0.575 | 46.14 | 3.61 | 1.83 | 0.403 | |
| Restart | 0.333 | 53.16 | 5.74 | 3.48 | 0.000 |
| Method | ||||||
|---|---|---|---|---|---|---|
| AquaMend (ours) | 80/80 | 0.625 / 1.875 | 0.250 / 0.250 | 2.125 | 0.9941 | 1.000 |
| Decision-theoretic troubleshooting | 80/80 | 0.750 / 2.250 | 0.075 / 0.075 | 2.325 | 0.9936 | 1.000 |
| Rollback-only | 20/80 | 1.000 / 3.000 | 0 / 0 | 14.750 | 0.9594 | 0 |
| Restart | 80/80 | 1.000 / 0 | 0 / 0 | 362.870 | 0 | 0 |
| Method | Realized cost | Expected policy risk | Gap to oracle | |
|---|---|---|---|---|
| Oracle (reference) | 4,000 | 7.406 | – | 0.000 |
| Exact Adaptive | 4,000 | 8.560 | 9.253 | 1.154 |
| Lookahead (depth 2) | 4,000 | 8.565 | 9.255 | 1.159 |
| AquaMend (myopic) | 4,000 | 8.594 | 9.290 | 1.188 |
| Troubleshooting | 4,000 | 8.599 | 9.557 | 1.193 |
| Decoupled | 4,000 | 11.008 | 11.756 | 3.602 |
| Ablation | Removed/ varied | Hypothesis | ||
|---|---|---|---|---|
| No re-probe | rollback only | re-probing is necessary to lower | 38.157 | 0.213 |
| No attribution | treat all as External | attribution avoids needless rollback | 5.625 | 0.900 |
| No source separation | self and external merged | cross-modal selection matters for drift | 5.541 | 0.901 |
| Detect obj. vs. VoI | swap criterion | detection needs a falsification objective | 5.431 / 10.335 | 0.901 / 0.708 |
| Chain vs. fine | graph granularity | chain benefit is zero rollback on false alarms | 6.830 / 5.360 | 0.999 / 0.999 |
| proxy | Indep / Beam / MC | proxy accuracy vs. cost | 5.431 / 5.431 / 5.432 | 0.901 / 0.901 / 0.901 |
| Method | Sensing | Rollback | Continue | Residual | Total |
|---|---|---|---|---|---|
| AquaMend | 0.463 | 2.847 | 53.672 | 14.375 | 71.356 |
| DTT | 0.450 | 14.436 | 47.543 | 9.688 | 72.116 |
| No extra re-probe | 0.325 | 4.434 | 45.191 | 34.375 | 84.326 |
| Full restart | 0.781 | 24.495 | 51.315 | 14.375 | 90.966 |
| Linear chain | 0.444 | 11.899 | 40.923 | 23.125 | 76.391 |
| Method | Mean-loss CI | ZRR | |||
|---|---|---|---|---|---|
| AquaMend | [64.833, 78.688] | 0.256 | 0.625 | 0.312 | 1.000 |
| DTT | [66.927, 77.555] | 0.185 | 0.563 | 1.469 | 1.000 |
| No extra re-probe | [78.399, 90.877] | 0.348 | 0.000 | 0.469 | 1.000 |
| Full restart | [83.443, 98.663] | 0.000 | 2.000 | 2.562 | 0.000 |
| Linear chain | [70.046, 83.426] | 0.305 | 0.688 | 1.281 | 1.000 |
| Stage | Scenes | AquaMend | All cand. | All cand. AquaMend [95% CI] |
|---|---|---|---|---|
| All | 32 | 2.074 | 2.364 | 0.290 [0.212, 0.360] |
| Early | 16 | 1.703 | 2.362 | 0.659 [0.499, 0.795] |
| Late | 16 | 2.445 | 2.366 | 0.080 [ 0.106, 0.051] |
| Method | Stage | Calls | Elig. | Used | Uncov. | Reduced | Enum. | Eval. | Branches |
|---|---|---|---|---|---|---|---|---|---|
| AquaMend | Early | 25 | 24 | 24 | 0 | 20 | 96 | 48 | 7383 |
| AquaMend | Late | 27 | 24 | 0 | 24 | 0 | 96 | 96 | 11094 |
| All cand. | Early | 25 | 0 | 0 | 0 | 0 | 96 | 96 | 11071 |
| All cand. | Late | 27 | 0 | 0 | 0 | 0 | 96 | 96 | 11094 |
| Chain | Early | 35 | 31 | 31 | 0 | 26 | 110 | 52 | 7692 |
| Chain | Late | 30 | 28 | 0 | 28 | 0 | 102 | 102 | 11220 |
| Scene | Residual declarations | Safe stop | Residual | Complete loss |
|---|---|---|---|---|
| s05 | Target | Yes | 80 | 94.450 |
| s13 | Target | Not recorded | 80 | 146.370 |
| s19 | Quantities A/B/C, target, sensing | Yes | 150 | 157.854 |
| s29 | Quantities A/B/C, target, sensing | Yes | 150 | 164.362 |
| Setting | Sensing | Rollback | Continue | Mean loss | Difference [95% CI] |
|---|---|---|---|---|---|
| Original | 0.475 | 2.828 | 50.611 | 63.914 | 0.000 [0.000, 0.000] |
| Concentrated | 0.475 | 2.828 | 50.611 | 63.914 | 0.000 [0.000, 0.000] |
| Dispersed | 0.525 | 2.863 | 50.711 | 64.099 | 0.186 [0.023, 0.441] |
| Negative bias | 0.475 | 2.828 | 50.611 | 63.914 | 0.000 [0.000, 0.000] |
| Positive bias | 0.550 | 3.889 | 51.525 | 65.964 | 2.051 [0.023, 5.851] |
| Event | Invalid rate | Mean | Brier | ECE |
|---|---|---|---|---|
| binding_A | 0.2500 | 0.2500 | 0.0000 | 0.0000 |
| binding_B | 0.2500 | 0.2500 | 0.0000 | 0.0000 |
| binding_C | 0 | 0.0000 | 0.0000 | 0.0000 |
| rank_A | 0.2500 | 0.2030 | 0.1143 | 0.1085 |
| rank_B | 0 | 0.0622 | 0.0310 | 0.0622 |
| rank_C | 0 | 0.0000 | 0.0000 | 0.0000 |
| Event | Invalid rate | Mean | Brier | ECE |
|---|---|---|---|---|
| binding_A | 0 | 0.0000 | 0.0000 | 0.0000 |
| binding_B | 0 | 0.0000 | 0.0000 | 0.0000 |
| binding_C | 0 | 0.0000 | 0.0000 | 0.0000 |
| rank_A | 0.1250 | 0.2205 | 0.1106 | 0.0955 |
| rank_B | 0.1250 | 0.0788 | 0.0228 | 0.0462 |
| rank_C | 0 | 0.0000 | 0.0000 | 0.0000 |
| Setting | Changed commands | Mean posterior TV | Mean predicted loss |
|---|---|---|---|
| Original | 0/32 | 0.0000 | 63.9249 |
| Concentrated | 1/32 | 0.0516 | 63.7192 |
| Dispersed | 8/32 | 0.0644 | 64.4001 |
| Negative bias | 0/32 | 0.1450 | 62.5118 |
| Positive bias | 8/32 | 0.1446 | 65.2785 |