ChunkTrust: Adapting Execution Horizons for Robot Policies with Action-Expert Evidence
Organizations: Tsinghua University · Beijing Academy of Artificial Intelligence (BAAI) · Shenzhen Technology University · Renmin University of China · Hefei University of Technology · Jiangnan University · Chongqing University · The Chinese University of Hong Kong
Abstract
Robot foundation policies predict action chunks, but how many actions to execute before replanning depends on the current task phase. We introduce ChunkTrust, which treats the execution horizon as a latent variable inferred from action-expert evidence rather than a fixed hyperparameter. Its training-free Action-aware Horizon Selector (AHS) combines intra-chunk spectral stability of generation traces with inter-chunk continuity between executed history and predicted actions. An online Beta posterior with kernel forgetting tracks horizon preferences across replans. A lightweight Query-based Horizon Adapter (QHA) optionally learns a context-conditioned dense prior from complementary evidence, fused with current evidence and episode-local Beta memory while the base policy remains frozen. Across RoboTwin2.0 and RoboCasa GR1 Tabletop, AHS improves overall task-averaged success for each evaluated base-policy configuration, including gains of +6.80 percentage points on over all 50 RoboTwin2.0 tasks and +9.67 percentage points on Qwen3GR00T in RoboCasa. AHS+QHA raises the gain over Base to +9.44 percentage points on the eight-task evaluation. On four real-world household tasks, AHS improves the equal-task mean normalized process score from 50.4% to 57.5%. Ablations examine the contributions of both evidence terms, temporal memory, and the learned prior. Project page is https://hf618.github.io/ChunkTrust.github.io/
Figures & tables
| Policy | Task Scope | Train Recipe | Base | +AHS | (pp) |
| RoboTwin2.0 Easy and Hard | |||||
| 50 tasks | multitask post-training | 56.70 | 63.50 | +6.80 | |
| 8 tasks | task-specific post-training | 18.19 | 21.38 | +3.19 | |
| 8 tasks | task-specific post-training | 29.63 | 36.25 | +6.63 | |
| Fast-WAM | 8 tasks | multitask post-training | 86.81 | 88.13 | +1.31 |
| RoboCasa GR1 Tabletop | |||||
| A. QHA training-task evaluation | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ( Black et al., 2024 ) | ( Intelligence et al., 2025 ) | |||||||||||
| Task | Base | AHS | AHS+QHA | Base | AHS | AHS+QHA | ||||||
| Easy | Hard | Easy | Hard | Easy | Hard | Easy | Hard | Easy | Hard | Easy | Hard | |
| Blocks Ranking RGB | 19.00 | 0.00 | 28.00 | 1.00 | 31.00 | 5.00 | 36.00 | 20.00 | 48.00 | 29.00 | 56.00 | 33.00 |
| Handover Block | 41.00 | 10.00 | 41.00 | 5.00 | 64.00 | 9.00 | 44.00 | 14.00 | 44.00 | 14.00 | 32.00 | 10.00 |
| Handover Mic | 100.00 | 2.00 | 100.00 | 24.00 | 100.00 | 30.00 | 98.00 | 64.00 | 100.00 | 56.00 | 98.00 | 59.00 |
| Candidate set | Success (%) | Success | Replan time (ms) | Overhead (%) | Replan counts | Ep. time (s) | |
|---|---|---|---|---|---|---|---|
| Fixed | – | 22.00 | – | 95.83 | – | 4.22 | 22.43 |
| 5 | 33.00 | +11.00 | 96.50 | 0.70 | 8.36 | 23.17 | |
| 10 | 32.50 | +10.50 | 96.84 | 1.05 | 9.34 | 23.08 | |
| 50 | 34.88 | +12.88 | 99.28 | 3.60 | 10.57 | 21.39 |
| Method | Easy (%) | Hard (%) | Overall (%) |
|---|---|---|---|
| Base (default) | 20.75 | 4.00 | 12.38 |
| Inter-chunk only | 23.50 | 4.00 | 13.75 |
| Intra-chunk only | 23.75 | 5.25 | 14.50 |
| Both evidence terms, no posterior | 25.50 | 3.75 | 14.63 |
| Full AHS | 24.25 | 5.75 | 15.00 |
| Full AHS+QHA | 27.50 | 4.00 | 15.75 |
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Task and sub-step | 0 | 0.5 | 1 |
|---|---|---|---|
| Bread: grasp | Fails to grasp | Multiple attempts | Smooth first attempt |
| Bread: handover | Handover fails | Unstable/awkward receiving grasp | Stable, aligned grasp |
| Bread: place on plate | Not placed | Poor alignment or rough placement | Clean placement |
| Drink: push | Only tilts, no useful displacement | Acceptable position, tilt/misalignment | Good position, stable alignment |
| Drink: grasp | Fails to grasp | Multiple attempts | Smooth first attempt |
| Drink: place in basket | Fails or drops drink | Rough placement/collision | Clean placement |
| Group | Intra risk | Inter risk | Episodes | Failures | Failure rate |
|---|---|---|---|---|---|
| Both low | Low | Low | 313 | 157 | 50.2% |
| Inter only | Low | High | 193 | 144 | 74.6% |
| Intra only | High | Low | 487 | 417 | 85.6% |
| Both high | High | High | 607 | 591 | 97.4% |
| Task | Setting | AHS | QHA-only | AHS+QHA | (pp) |
|---|---|---|---|---|---|
| Blocks Ranking RGB | Easy | 44.00 | 48.00 | 57.00 | +13.00 |
| Blocks Ranking RGB | Hard | 33.00 | 29.00 | 30.00 | -3.00 |
| Place Bread Basket | Easy | 52.00 | 50.00 | 49.00 | -3.00 |
| Place Bread Basket | Hard | 31.00 | 31.00 | 35.00 | +4.00 |
| Overall | Both | 40.00 | 39.50 | 42.75 | +2.75 |
| Task | Easy | Hard | (pp) | ||
|---|---|---|---|---|---|
| Base | AHS | Base | AHS | ||
| adjust bottle | 100.00 | 100.00 | 90.00 | 95.00 | +2.50 |
| beat block hammer | 70.00 | 80.00 | 15.00 | 50.00 | +22.50 |
| blocks ranking rgb | 80.00 | 95.00 | 40.00 | 70.00 | +22.50 |
| blocks ranking size | 25.00 | 45.00 | 25.00 | 35.00 | +15.00 |
| click alarmclock | 60.00 | 65.00 | 45.00 | 45.00 | +2.50 |
| Fast-WAM ( Yuan et al., 2026 ) | X-VLA ( Zheng et al., 2026 ) | |||||||
| Task | Base | AHS | Base | AHS | ||||
| Easy | Hard | Easy | Hard | Easy | Hard | Easy | Hard | |
| Blocks Ranking RGB | 99.00 | 98.00 | 100.00 | 100.00 | 87.00 | 34.00 | 95.00 | 38.00 |
| Handover Block | 94.00 | 81.00 | 91.00 | 82.00 | 89.00 | 2.00 | 88.00 | 1.00 |
| Handover Mic | 100.00 | 99.00 | 100.00 | 100.00 | 97.00 | 1.00 | 100.00 | 1.00 |
| Hanging Mug | 66.00 | 64.00 | 71.00 | 69.00 | 32.00 | 8.00 | 35.00 | 4.00 |
| Task | QwenFAST +Qwen3VL | QwenPI +Qwen3VL | Isaac-GR00T N1.5 | Isaac-GR00T N1.6 | Qwen3GR00T +Qwen3VL | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| Base | AHS | Base | AHS | Base | AHS | Base | AHS | |||
| PnPBottleToCabinetClose | 38.0 | 26.0 | 64.0 | 68.0 (+4.0) | 51.5 | 54.0 (+2.5) | 46.0 | 64.0 (+18.0) | 64.00 | 68.00 (+4.00) |
| PnPCanToDrawerClose | 44.0 | 62.0 | 18.0 | 12.0 (-6.0) | 13.0 | 12.0 (-1.0) | 80.0 | 80.0 (+0.0) | 58.00 | 56.00 (-2.00) |
| PnPCupToDrawerClose | 56.0 | 42.0 | 12.0 | 4.0 (-8.0) | 8.5 | 14.0 (+5.5) | 54.0 | 52.0 (-2.0) | 34.00 | 40.00 (+6.00) |
| PnPMilkToMicrowaveClose | 44.0 | 50.0 | 38.0 | 34.0 (-4.0) | 14.0 | 20.0 (+6.0) | 48.0 | 42.0 (-6.0) | 40.00 | 44.00 (+4.00) |
| PnPPotatoToMicrowaveClose | 14.0 | 42.0 | 54.0 | 36.0 (-18.0) | 41.5 | 50.0 (+8.5) | 28.0 | 28.0 (+0.0) | 22.00 | 30.00 (+8.00) |
| RoboTwin2.0 | |||||
|---|---|---|---|---|---|
| Selector | SR (%) | Calls/ep | Policy infer | Total infer | Wall |
| Global fixed ( ) | 14.84 | 25.96 | 2.75 | 2.75 | 54.33 |
| Jerk-min | 14.06 | 21.14 | 2.28 | 2.28 | 52.01 |
| EMA | 10.94 | 19.72 | 2.05 | 2.07 | 46.72 |
| AHS (Full Beta) | 15.63 | 20.71 | 2.26 | 2.28 | 53.47 |
| Instantaneous (reference) | 13.28 | 19.58 | — | — | — |
| Calls/ | Policy inference | Episode wall | |
| Method | episode | (s/episode) | time (s/episode) |
| A. Full 50-task evaluation 2,000 episodes per method | |||
| Base | 8.00 | 0.97 | 38.19 |
| AHS | 14.50 | 1.55 | 36.83 |
| B. QHA timing evaluation 2 held-out tasks, 400 episodes per method | |||
| AHS | 33.77 | 3.29 | 68.60 |
| Method | SR (%) | Wait (%) | Wait s/ep | Calls/ep | Infer s/ep | Mean age (s) | ||
|---|---|---|---|---|---|---|---|---|
| +0 ms | +100 ms | +200 ms | ||||||
| All four tasks | ||||||||
| Sync-Fixed | 19.50 | 19.50 | 19.25 | 3.77 | 4.432 | 12.89 | 5.04 | 5.02 |
| Sync-AHS | 26.25 | 24.25 | 23.75 | 6.80 | 7.893 | 22.95 | 8.59 | 3.00 |
| RTC-Fixed | 23.00 | 21.50 | 22.00 | 0.31 | 0.344 | 13.32 | 6.47 | 4.92 |
| RTC-AHS | 26.50 | 25.75 | 26.75 | 0.31 | 0.344 | 24.30 | 10.62 | 3.07 |
| Method | SR (%) | Wait (%) | Wait s/ep | Calls/ep | Infer s/ep | Mean age (s) | ||
|---|---|---|---|---|---|---|---|---|
| +0 ms | +100 ms | +200 ms | ||||||
| All four tasks | ||||||||
| Sync-Fixed | 38.00 | 35.50 | 37.00 | 3.90 | 3.856 | 11.21 | 4.38 | 5.12 |
| Sync-AHS | 46.00 | 46.25 | 44.75 | 6.63 | 6.310 | 18.35 | 6.90 | 3.16 |
| RTC-Fixed | 38.75 | 40.25 | 39.00 | 0.37 | 0.344 | 11.50 | 5.57 | 5.00 |
| RTC-AHS | 46.00 | 48.50 | 42.75 | 0.37 | 0.344 | 20.32 | 8.80 | 3.18 |
| Delay (ms) | Method | Wait (%) | Wait s/ep | Calls/ep | Infer s/ep | Mean age (s) |
|---|---|---|---|---|---|---|
| Easy | ||||||
| 0 | Sync-Fixed | 1.67 | 1.600 | 11.11 | 4.32 | 4.93 |
| Sync-AHS | 2.87 | 2.614 | 18.15 | 6.78 | 3.01 | |
| RTC-Fixed | 0.16 | 0.147 | 11.20 | 5.38 | 5.00 | |
| RTC-AHS | 0.16 | 0.147 | 17.86 | 7.81 | 3.24 | |
| 100 | Sync-Fixed | 2.82 | 2.755 | 11.29 | 4.37 | 4.98 |
| Parameter | Default | Tested | Easy | Hard | Overall |
|---|---|---|---|---|---|
| Forgetting | 0.99 | 0.95 | 46.88 | 26.75 | 36.81 |
| 0.995 | 46.50 | 25.88 | 36.19 | ||
| Update strength | 1 | 0.5 | 45.38 | 24.38 | 34.88 |
| 2 | 47.38 | 26.88 | 37.13 | ||
| Kernel bandwidth | 10 | 5 | 45.50 | 26.13 | 35.81 |
| 20 | 44.25 | 24.50 | 34.38 |