Agents are deployed in applications from trip planners and stock trading to IT incident triage. In most cases, LLM agents work autonomously with minimal rule-based safeguarding, leading to cost and safety issues from irreversible actions. Recent works resolve this either by using a safeguard agent to monitor behavior or evaluating logs post-hoc. The first adds cost and latency to every step; the second delivers its verdict after the run, when tokens are burned and damage is done. To overcome this, we propose OnTrack, a streaming monitoring mechanism that compares an agent's steps and dependencies against recorded successful runs to alert users or block the agent in about a millisecond per step. We study this problem in three regimes of decreasing access: full reference access (historical runs and tool schemas), intermediate access (only tool schemas), and no prior knowledge (only step logs as generated). Expectation of OnTrack's monitoring capabilities reduces as data access drops, ranging from plan violation detection to identifying loops, stalls, and repeated tool calls. Finally, we evaluate OnTrack using SWE-bench trajectories. Based on the first 8 steps, our method ranks failing trajectories below succeeding ones better than content similarity approaches (+0.057 AUROC). With an abort policy, we save about 18% of compute that would be burned on failing runs, where 83% of interrupted runs were actually heading to failure (5 out of 6 aborts were correct).
Figures & tables
Figure 1: The monitor’s inputs. Events arrive one at a time and grow the prefix graph Pt (dashed: a dependency discovered retroactively). References R1,…,RK are known-good executions in the same form. Each graph carries a matrix of pairwise dependency-hop distances ( DP grows row by row; DRk is precomputed), and the alignment T of Sec. 4 matches the two sides step by step.
Figure 2: Information regimes. Each capability band runs for as long as its inputs exist and terminates when they are removed: without references there is no plan to deviate from (L2 ends); without schemas there is nothing to gate (L3 ends); the vital signs of L1 never turn off. Bands are stacked by how long they survive (bedrock at the bottom), not by layer number; L1–L3 are the monitor’s three layers, defined in the architecture paragraph in Sec. 4 . The strip underneath shows the orthogonal quality axis: the richer the provenance of the references, the more authority a verdict carries (Sec. 4.6 ).
Figure 3: The two ingredients on a toy example: three agent steps against a four-step reference. In the cost table C , dark means implausible: “grep error” matches search cheaply, “edit migration” matches edit , and the hallucinated “fetch weather” is expensive against every reference step. Minimizing Eq. ( 2 ) yields the alignment T (dark = matched mass): the two valid steps place their mass on their counterparts, while the hallucinated step’s row stays nearly empty — the objective preferred paying the unmatched-mass penalty over forcing a bad match. That empty row is the leakage signal of Eq. ( 4 ).
Signal
Definition / role
node leakage ℓt
Eq. ( 4 ) on the newest node; hallucination / digression / prematurity, routed through disambiguation
match cost cˉt
∑jTtjMtj/(∑jTtj+ϵden) ; “matched, but expensively” (wrong tool on the right step); the ϵden floor covers fully-destroyed rows (distinct from the entropic ε )
coverage velocity
ρt−ρt−w ; stall = activity without progress
loop score (L1)
near-duplicate predecessor (action, tool, arguments), with a retry exemption : duplicates of an errored step are exempt up to rmax attempts
info gain (L1)
ΔIt=1−maxi<tcos(ϕ(et),ϕ(ei)) over effect embeddings; spinning = new calls, no new state
strategy switch
the soft-min argmink changes; benign once, oscillation warns
Table 1: Per-event signals. All are O(t) or free by-products of the transport update.
Scenario
Expected
Fired
Lat.
Layer
A clean in-order run
silence
silence ✓
—
—
B valid reordering
silence
silence ✓
—
L2
C strict loop (2-step cycle)
LOOP
LOOP
0
L1
D hallucination streak
OFF-TRK
OFF-TRK
2
L2
E causal inversion
C-INV
C-INV
0
L2
F retry after tool error
silence
silence ✓
—
L1
Table 2: Controlled fault-injection results ( OFF-TRK = OFF_TRACK , C-INV = CAUSAL_INVERSION ). Latency in events after anomaly onset (target ≤2 ); the last column attributes each behavior to the responsible layer. Scenario B finalizes to the in-order score exactly ( 0.913=0.913 ), the empirical face of Prop. 2 . In D, the two grace-window events preceding escalation surface as EXPLORING , not silence.
Scenario
Expected
Result
G-a
valid irrev., prereqs + ref
allow
allow ✓
G-b
missing prereq artifact
block
block ✓
G-c
schema-valid, no ref match
policy
block (conserv.)
G-d
no schema, no references
policy
allow (no check)
False blocks, valid irrev. calls
0/50 (95% bound ≈ 6%)
Table 3: L3 gate tests — a small validation suite, not a reliability claim. G-c/G-d record implemented policy: with references present the gate requires both the schema check and the proposal match, checked against the current best (lowest-loss) reference only (conservative); in the absence of both schemas and references the gate operates allow-by-default — an explicit deployment policy, not a safety guarantee.
Variant
Loop
Halluc
Invert
Stall
FPR
Full system
30/30
12/30
12/30
30/30
13/40
True no-mask ( ν(t)=ν )
30/30
13/30
12/30
30/30
14/40
No GW ( θ=0 )
30/30
21/30
25/30
30/30
29/40
L1-only (no refs)
30/30
18/30
19/30
30/30
31/40
Prereq-rule baseline
30/30
26/30
24/30
30/30
32/40
Frozen frontier ( κ=100 )
30/30
30/30
26/30
30/30
38/40
Table 4: Component study on real SWE-agent traces at the shipped verdict thresholds, reported as raw counts (detected/injected per fault; false-positive traces over n=40 benign traces; trace-level ≥ 1-severe-flag criterion — the deployed abort policy thresholds flag density instead, Table 6 , hence its far lower false-stop rate). Wilson 95% intervals at these sample sizes span roughly ±15 percentage points (e.g., 12/30 = 40% [24.6%,57.7%] ; 13/40 = 32.5% [20.1%,48.0%] ), so only the large contrasts are meaningful. “True no-mask” removes the frontier mask entirely; “frozen frontier” ( κ=100 , nothing ever satisfied) is a mis-set mask, included to show that a wrongly configured mask is worse than none.
k
OnTrack
LinUOT
Cosine
n_st.
GED
Δ vs. cosine
5
.612
.595
.552
.622
.574
+.060[.037,.083]
8
.631
.628
.574
.613
.576
+.057[.033,.080]
10
.622
.614
.589
.598
.552
+.033[.007,.057]
15
.635
.626
.633
.580
.551
+.002[−.022,.027]
20
.648
.638
.646
.554
.566
+.001[−.029,.031]
Table 5: Primary protocol: AUROC at partial trajectory (first k steps), every evaluation-partition trace with ≥k steps ( N=2,055 at k=5 down to 1,405 at k=20 ), all methods scored on the same traces, instance-level bootstrap 95% CIs ( n=1,000 ) on the OnTrack − cosine difference. The Δ CI excludes zero for k≤10 .
Threshold
True stop
False stop
Cost saved
Prec. (dep.)
0.20
33.3%
29.3%
28.0%
83%
0.40
27.3%
26.7%
20.9%
81%
0.60
23.3%
20.7%
17.9%
83%
0.70
20.7%
19.3%
15.6%
82%
Table 6: Abort threshold sweep on severe flag density ( n=150 per class, balanced). Precision is reported at the corpus base rate (81% of traces unresolved), which is the deployment-relevant figure; on the balanced sample it is necessarily near 50%.
Method
AUROC
OnTrack score
0.979
OnTrack flags
0.977
n_steps
0.978
Markov surprise
0.905
OnTrack leak
0.958
Cosine
0.557
Table 7: Exit-status detection: cost-limit vs. clean submission (balanced 50+50 sample). OnTrack matches n_steps ; the value added is the calibrated mid-trace trigger (Table 6 ), not the full-trace AUROC.
Signal
AUROC
OnTrack score
0.581
OnTrack leak
0.550
OnTrack flags
0.577
Cosine
0.643
n_steps
0.576
Markov surprise
0.573
Table 8: Resolved vs. unresolved (50+50 balanced), one row per trajectory-level signal. Bootstrap 95% CI: OnTrack score [0.466,0.689] , cosine [0.531,0.759] — overlapping.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
k
N
OnTrack
Cosine
Δ
Seeds +
5
200
.618±.047
.566±.057
+.052±.045
4/5
8
193
.649±.036
.581±.032
+.067±.028
5/5
10
184
.649±.020
.588±.022
+.061±.022
5/5
15
155
.656±.032
.629±.007
+.027±.027
4/5
20
128
.667±.010
.653±.022
+.013±.022
3/5
Appendix
Table 9: Secondary (balanced multi-seed) protocol: mean ± std over 5 seeds; Seeds + counts seeds with Δ>0 . Absolute AUROC runs a few points higher than the primary protocol (balanced samples over-weight the resolved class relative to the 81%-unresolved population) and the deltas are somewhat larger; the direction and the early- k concentration of the advantage agree with Table 5 , whose numbers we cite throughout.