Generating a plausible robot action does not establish that current observations justify its execution. Motivated by exploratory observations of high-confidence visual outputs under severe occlusion, we present PIER, an execution-authorization interface that separates evidence checks, decision provenance, and stage-scoped re-observation from hardware control. The deterministic gate evaluates declared visual and tactile inputs, while its caller maintains a budget of at most one re-observation per stage. We evaluate the implementation using 1,600 threshold-grid cases and 1,200 paired synthetic traces spanning score noise, missing tactile inputs, stale observations, and falsely reassuring scores. The finite grid yields zero declared invariant violations, and a matched Boolean baseline reproduces all non-recovery decisions. Under synthetic score noise, re-observation reduces valid-state denials from 57/120 to 18/120 while increasing invalid-state proceeds from 5/120 to 7/120. Stale and falsely reassuring inputs expose limitations that threshold checks alone cannot resolve. Exploratory visual, tactile, and robot setup records provide context but do not establish physical task performance. These results characterize an inspectable authorization interface and its input-contract limitations, without claiming superiority over equivalent rule logic, calibrated tactile accuracy, or certified physical safety.
Figures & tables
Fig. 1: PIER recovery-mode authorization and external execution boundary. Source timestamps do not imply enforced freshness or cross-sensor alignment. Re-observation is a software request, limited to at most once per stable stage ID within one session; the caller maintains the count. The record associates evidence, stage, and decision reason; software revision is package-level provenance. Proceed does not bypass adapter checks or operator approval.
TABLE II: Rerun initial fault responses. P/R/S: proceed/re-observe/stop. These checks validate declared logic, not sensor accuracy.
Field
Event 3
Event 4
Timestamp (ns)
1,040,000,000
1,060,000,000
Stage
route
route
Evidence IDs
rgb-003, tac-003
rgb-004, tac-004
(v,c,s)
(.90,.91,.82)
(.91,.92,.07)
Retries before / after
0 / 1
1 / 1
Decision
RETRY
PROCEED
TABLE III: Concrete trace case: two records from the same route stage. IDs identify synthetic evidence; no sensor files or robot motion are implied.
Condition
Vision only
V–T
V–T + R
F + V–T + R
Clean
120 / 0
0 / 0
0 / 0
0 / 0
Gaussian score noise
97 / 31
5 / 57
7 / 18
7 / 18
Tactile dropout
120 / 0
0 / 42
0 / 18
0 / 18
Stale reassuring input
120 / 0
120 / 0
120 / 0
0 / 120
Fresh falsely reassuring input
120 / 0
120 / 0
120 / 0
120 / 0
TABLE IV: Paired synthetic stress test. Each cell is invalid-state proceeds / valid-state denials, with separate denominators of 120 each. R: one re-observation; F: experimental freshness wrapper. All policies share inputs. No rates are interpreted as deployment risk.
Evidence
Scope and provenance
Supported interpretation and limit
Finite gate grid
1,600 cases; freshly rerun
Declared implementation invariants; not universal or physical safety.
Synthetic stress
1,200 paired traces; 6,000 policy evaluations; new
Noise/retry trade-off, stale-input and false-confidence limits; no empirical sensor model.
Contract fixture
Nine tests; one synthetic bundle; rerun
Cross-file and artifact integrity checks; no physical truth or general schema-compliance guarantee.
Fault/trace checks
20 mode–scenario cases; five trace events; rerun
Initial decision semantics and stage-level retry accounting; no robot replay.
Closed-loop software
60 executions / 20 matched seeds; reported
Historical context only; simulator and episode outputs unavailable.
D405/SAM3 pilot
Nine observations; exploratory summary
Motivating observation only; rule configuration and raw outputs unavailable.
TABLE V: Consolidated evidence ledger. Different units and evidence classes are never pooled as robot trials. “Reported” indicates retained manuscript/transcript results, not a fresh raw-data rerun.
Fig. 2: Exploratory visual pilot: retained clear, partial-occlusion, and severe-occlusion examples. The image annotations report a downstream visual rule proceeding in all three, including severe occlusion with confidence 0.938 and an expected stop label. Numerical annotations are transcribed in Table VI for readability. Full configuration and primary outputs are unavailable; this is not a reproduced inference benchmark or evidence of tactile correction.
Condition
Clear
Partial
Severe
Maximum confidence
0.898
0.898
0.938
Union visual support
0.0057
0.0064
0.0040
Rule decision
Proceed
Proceed
Proceed
Expected decision
Proceed
Proceed
Stop
Agreement
Correct
Correct
Incorrect
TABLE VI: Representative SAM3 comparison transcribed from Fig. 2 . Confidence and support values belong to the displayed examples, not condition averages.
Prompt phase
Frames
Median MAD
No contact (full phase)
300
2.6522
Mixed touch/release
660
2.7168
Release
241
2.7229
TABLE VII: Reported Xense phase statistics. Prompt phases are not contact ground truth; frames are temporally correlated.
Fig. 3: Exploratory physical observation panel: D405 RGB/geometric views, Xense views, and the recorded MAD timeline. The original full-record and initialization-aware views are preserved without reconstructing raw samples. Shaded bands denote operator prompts, not independently verified contact states.
Fig. 4: WidowX with the modified Xense fixed tool. This setup photograph documents hardware context; it does not establish tool-to-hole alignment, calibrated geometry, or a successful insertion.
Attempt
Reported outcome
001
Preflight SAFE_STOP; zero Cartesian dispatches.
002
One 5-mm dispatch; maximum signed progress 0.190 mm; J2 effort delta 1.384 Nm exceeds 1.350 Nm; SAFE_STOP.
Both
Zero gripper position/non-idle-mode commands; zero automatic returns or retries.
TABLE VIII: Exploratory fixed-tool qualification. Values are retained summary claims; the original JSON records were not recovered for independent verification.
Demonstration-conditioned policies provide a natural interface for specifying robot behavior, yet long-horizon manipulation remains difficult when visually similar states recur across different stages or when demonstrations and executions proceed at different speeds. We identify the resulting failure mode as stage confusion and introduce Progress-Aligned Context for Execution (PACE), a stateful method that continually reinterprets a complete demonstration according to realized execution progress. PACE compresses the demonstration into ordered multimodal prompt tokens and uses training-only dual-edge attention supervision to expose its latent stage structure. During execution, an episode-local fast-weight memory causally encodes realized action-observation transitions and modulates prompt cross-attention, producing a progress-aligned context for a unified diffusion action expert without test-time stage labels or stage-specific policies. PACE improves success from 88.9% to 94.0% on LIBERO-Gen Goal Chain, from 79.1% to 83.3% on Spatial Combination, and from 33.3% to 73.3% on the two-step Block Routing tasks. Failure analysis further indicates that structured demonstration alignment and causal execution memory jointly mitigate stage confusion.
Yenan Chen, Junjie Shi, Lu Chen +2
College of Control Science and Engineering, Zhejiang University, Hangzhou 310027, China
Although Vision--Action (VA) and Vision--Language--Action (VLA) policies have advanced robotic manipulation, their evaluation remains dominated by binary success rates, which obscure process-level differences among executions that complete the same task. We introduce Eval-Actions, a diagnostic evaluation methodology and real-robot benchmark for fine-grained execution-quality assessment of learned manipulation policies. Eval-Actions combines criteria-based Expert Grading (EG), Rank-Guided (RG) labels that align measurable motion indicators with expert rankings, and Chain-of-Thought-style (CoT) annotations that explain observable quality differences. The benchmark contains 13K+ teleoperated and policy-generated real-robot episodes covering 150+ tasks and approximately 52 hours of recordings with RGB-D videos, robot-state trajectories, task descriptions, and success/failure labels. Its densely annotated subset provides EG/RG/CoT supervision for training and evaluation. We further provide AutoEval, a reference multimodal evaluator that predicts quality scores, task outcomes, and diagnostic explanations from RGB temporal evidence and compact kinematic summaries. On the annotated Eval-Actions test split, AutoEval-S achieves Spearman rank correlations (SRCCs) of 0.81 and 0.84 under EG and RG, with success detection accuracies of 90.6% and 91.0%; AutoEval-P reaches 0.70 SRCC under CoT. Analyses of expert consistency, physical-metric baselines, modality ablations, structured generalization, and offline policy ranking show that Eval-Actions provides standardized, interpretable diagnostic signals complementary to success-rate evaluation.
Recent vision-language-action and diffusion-based robot policies often use action chunking, where each policy query predicts a sequence of future actions and the robot executes an open-loop prefix before re-querying. While this interface improves local motion continuity, deployment still requires choosing the execution horizon: how much of each predicted chunk should be executed before acquiring a new observation. However, our experiments show that success is strongly task-dependent and non-monotonic with respect to the execution horizon, making a single constant horizon an unreliable deployment rule. We propose PACE (Phase-Aware Chunk Execution), a training-free test-time execution method that selects the execution horizon online from the predicted chunk itself. PACE exploits the phase-dependent kinematic structure of manipulation trajectories by identifying low-speed transition points in the predicted speed profile and using them as candidate replanning boundaries. Because PACE uses only the predicted action chunk, it is plug-and-play and requires no retraining or access to policy internals. We validate PACE through large-scale evaluations in both simulation and real-robot settings. On 50 RoboTwin2.0 tasks, PACE raises the average success rate from 57.8% to 64.2%. In real-robot experiments on bimanual ALOHA and single-arm Franka platforms, PACE improves the average task score from 60.7 to 77.7 and the average success rate from 50.7% to 70.4%. Ablations and rollout-level analyses show that PACE adapts execution horizons across manipulation phases, shortening near transitions while preserving longer execution during coherent motion.
Junnan Nie, Jiayi Li, Chenghao Liu +5
Peking University · Peking University. · JD Explore Academy. +1