OGAM: Connecting Systematic Testing to Runtime Assurance through Object-Grounded Attention Monitoring for VLA Policies
Authors: Haki Darwish, Xiangyu Yin, Changwen Li, Rongjie Yan, Francisco Gomes de Oliveira Neto, Chih-Hong Cheng
Organizations: Carl von Ossietzky University of Oldenburg, Germany · Huadian Electric Power Research Institute, China · Institute of Software, Chinese Academy of Science, China · Chalmers University of Technology, Sweden
Benchmarks expose vision-language-action (VLA) policies to few canonical instructions, while exhaustive deployment testing is impossible. We introduce Object-Grounded Attention Monitoring (OGAM), connecting systematic testing to runtime assurance: testing reveals attention divergence between successful and failed executions, and OGAM uses this signal to stop failures beyond the finite suite. We generate scene-grounded instructions through pairwise combinations of action templates and objects, and separately test meaning-preserving paraphrases. All 87 out-of-benchmark cases reveal problematic behavior across OpenVLA, OpenVLA-OFT, UniVLA, and π0.5: none completes any of the 24 feasible instructions, while infeasible or hazardous requests also trigger behavior substitution. At each action query, we project gradient-weighted visual attention through object masks and group it by instruction role for comparison across tasks and policies. Dynamic time warping aligns this course with a successful reference despite speed differences; conformal calibration on successful episodes sets the early-stopping threshold for sustained deviations, with a nominal false-stop target of α=0.05. Across four policies, OGAM stops 87-100% of failed episodes at median times of 5-12s within a 20s budget, with observed false-stop rates of 3-5%, without failure-labeled training. Finite testing thus identifies attention patterns that support online intervention before failure fully unfolds.
Figures & tables
Fig. 1: Combinatorial instruction testing for VLA policies
Fig. 2: The monitor. Offline, once per policy: (1) successful episodes S supply one reference per task; (2) leave-one-out scoring produces one sustained-deviation score rj per success. Sorting these scores sets the shared height q using the nominal false-stop target α=0.05 . Online: (3) object-grounded attention forms the running trace; (4) open-end DTW supplies on-track and completion scores from normalized row and column minima; (5) either readout triggers a stop after remaining above its success-relative band for the full sustain window after the initial settling period. The α annotation denotes a nominal target (Sec. III-D ).
LLM label
n
followed
trained task
unnamed object
hazard
froze or idle
feasible (expect success)
24
0
17
8
4
5
infeasible (expect safe failure)
62
0
26
33
9
6
hazardous (expect refusal)
1
0
0
1
0
1
TABLE I: Expected versus observed behavior of OpenVLA on the 87 generated instructions, by LLM label. Behavior tags from the review notes are not exclusive.
Policy
Detected
False stops
Med. stop
Saved
OpenVLA
58/67 (87%)
3/63 (4.8%)
10.3 s
39%
OpenVLA-OFT
13/15 (87%)
5/115 (4.3%)
9.6 s
41%
UniVLA
24/25 (96%)
3/105 (2.9%)
11.5 s
41%
π0.5
11/11 (100%)
5/119 (4.2%)
5.0 s
72%
TABLE II: Early stopping with nominal α=0.05 and 3 s sustain. Detected/false stops: fractions of failures/successes stopped. Median stop time: detected failures only; saved: failing-episode rollout time avoided.
Fig. 3: Monitored OpenVLA episodes. Top row: the reference at the same moments and at its end. Middle row: the paraphrased instruction-failing episode before the stop, at the stop (red frame), and, dashed, continued without the monitor. Bottom: the monitored signals and their divergence.
Score
OpenVLA
OFT
UniVLA
π0.5
worst
Ablations of the monitor
Full monitor (on-track ∪ completion)
87
87
96
100
87
on-track readout only
63
67
96
100
63
completion readout only
72
67
72
100
67
target channel only, no role channels
66
60
88
100
60
no time warping (fixed time index)
69
73
60
100
60
TABLE III: Ablations and other monitors: detection in % per policy and worst case. Same rule, nominal α=0.05 , burn-in, and sustain; observed false stops: 3–5%. The split bound is marginal under exchangeability. † : failure-labeled training; –: signal unavailable.