Generating a plausible robot action does not establish that current observations justify its execution. Motivated by exploratory observations of high-confidence visual outputs under severe occlusion, we present PIER, an execution-authorization interface that separates evidence checks, decision provenance, and stage-scoped re-observation from hardware control. The deterministic gate evaluates declared visual and tactile inputs, while its caller maintains a budget of at most one re-observation per stage. We evaluate the implementation using 1,600 threshold-grid cases and 1,200 paired synthetic traces spanning score noise, missing tactile inputs, stale observations, and falsely reassuring scores. The finite grid yields zero declared invariant violations, and a matched Boolean baseline reproduces all non-recovery decisions. Under synthetic score noise, re-observation reduces valid-state denials from 57/120 to 18/120 while increasing invalid-state proceeds from 5/120 to 7/120. Stale and falsely reassuring inputs expose limitations that threshold checks alone cannot resolve. Exploratory visual, tactile, and robot setup records provide context but do not establish physical task performance. These results characterize an inspectable authorization interface and its input-contract limitations, without claiming superiority over equivalent rule logic, calibrated tactile accuracy, or certified physical safety.
Figures & tables
Fig. 1: PIER recovery-mode authorization and external execution boundary. Source timestamps do not imply enforced freshness or cross-sensor alignment. Re-observation is a software request, limited to at most once per stable stage ID within one session; the caller maintains the count. The record associates evidence, stage, and decision reason; software revision is package-level provenance. Proceed does not bypass adapter checks or operator approval.
TABLE II: Rerun initial fault responses. P/R/S: proceed/re-observe/stop. These checks validate declared logic, not sensor accuracy.
Field
Event 3
Event 4
Timestamp (ns)
1,040,000,000
1,060,000,000
Stage
route
route
Evidence IDs
rgb-003, tac-003
rgb-004, tac-004
(v,c,s)
(.90,.91,.82)
(.91,.92,.07)
Retries before / after
0 / 1
1 / 1
Decision
RETRY
PROCEED
TABLE III: Concrete trace case: two records from the same route stage. IDs identify synthetic evidence; no sensor files or robot motion are implied.
Condition
Vision only
V–T
V–T + R
F + V–T + R
Clean
120 / 0
0 / 0
0 / 0
0 / 0
Gaussian score noise
97 / 31
5 / 57
7 / 18
7 / 18
Tactile dropout
120 / 0
0 / 42
0 / 18
0 / 18
Stale reassuring input
120 / 0
120 / 0
120 / 0
0 / 120
Fresh falsely reassuring input
120 / 0
120 / 0
120 / 0
120 / 0
TABLE IV: Paired synthetic stress test. Each cell is invalid-state proceeds / valid-state denials, with separate denominators of 120 each. R: one re-observation; F: experimental freshness wrapper. All policies share inputs. No rates are interpreted as deployment risk.
Evidence
Scope and provenance
Supported interpretation and limit
Finite gate grid
1,600 cases; freshly rerun
Declared implementation invariants; not universal or physical safety.
Synthetic stress
1,200 paired traces; 6,000 policy evaluations; new
Noise/retry trade-off, stale-input and false-confidence limits; no empirical sensor model.
Contract fixture
Nine tests; one synthetic bundle; rerun
Cross-file and artifact integrity checks; no physical truth or general schema-compliance guarantee.
Fault/trace checks
20 mode–scenario cases; five trace events; rerun
Initial decision semantics and stage-level retry accounting; no robot replay.
Closed-loop software
60 executions / 20 matched seeds; reported
Historical context only; simulator and episode outputs unavailable.
D405/SAM3 pilot
Nine observations; exploratory summary
Motivating observation only; rule configuration and raw outputs unavailable.
TABLE V: Consolidated evidence ledger. Different units and evidence classes are never pooled as robot trials. “Reported” indicates retained manuscript/transcript results, not a fresh raw-data rerun.
Fig. 2: Exploratory visual pilot: retained clear, partial-occlusion, and severe-occlusion examples. The image annotations report a downstream visual rule proceeding in all three, including severe occlusion with confidence 0.938 and an expected stop label. Numerical annotations are transcribed in Table VI for readability. Full configuration and primary outputs are unavailable; this is not a reproduced inference benchmark or evidence of tactile correction.
Condition
Clear
Partial
Severe
Maximum confidence
0.898
0.898
0.938
Union visual support
0.0057
0.0064
0.0040
Rule decision
Proceed
Proceed
Proceed
Expected decision
Proceed
Proceed
Stop
Agreement
Correct
Correct
Incorrect
TABLE VI: Representative SAM3 comparison transcribed from Fig. 2 . Confidence and support values belong to the displayed examples, not condition averages.
Prompt phase
Frames
Median MAD
No contact (full phase)
300
2.6522
Mixed touch/release
660
2.7168
Release
241
2.7229
TABLE VII: Reported Xense phase statistics. Prompt phases are not contact ground truth; frames are temporally correlated.
Fig. 3: Exploratory physical observation panel: D405 RGB/geometric views, Xense views, and the recorded MAD timeline. The original full-record and initialization-aware views are preserved without reconstructing raw samples. Shaded bands denote operator prompts, not independently verified contact states.
Fig. 4: WidowX with the modified Xense fixed tool. This setup photograph documents hardware context; it does not establish tool-to-hole alignment, calibrated geometry, or a successful insertion.
Attempt
Reported outcome
001
Preflight SAFE_STOP; zero Cartesian dispatches.
002
One 5-mm dispatch; maximum signed progress 0.190 mm; J2 effort delta 1.384 Nm exceeds 1.350 Nm; SAFE_STOP.
Both
Zero gripper position/non-idle-mode commands; zero automatic returns or retries.
TABLE VIII: Exploratory fixed-tool qualification. Values are retained summary claims; the original JSON records were not recovered for independent verification.