Forensic reconstruction of LLM-agent actions requires not only recovering the correct value, but establishing which preserved record supports that finding. Tool logs, generated explanations, and local citation identifiers capture different parts of this evidence, yet a citation identifier does not establish a source unless its binding to a record is preserved. We audit this distinction using 64 mechanically checkable cases from saved AgentDojo Banking executions. Two LLM readers reconstruct source relationships under controlled variations in visible evidence and identifier-to-record bindings. We separately evaluate complete-record agreement, evidence-grounded findings, justified abstention, and unsupported assertions. With original identifiers and no binding table, Sonnet recovered every literal source location but made unsupported citation-source assertions in 26 of 28 cases requiring the missing relation; 22 nevertheless matched the complete reference. Adding explicit bindings improved grounded reconstruction for both readers, whereas identifier renaming alone provided no consistent remedy. A deterministic same-packet comparator correctly resolved the bounded task or abstained throughout. These results show that factual agreement alone is insufficient for evaluating forensic reconstruction of agent logs and motivate preserving explicit record bindings to distinguish supported findings from correct guesses.
Figures & tables
Figure 1: Both studies reuse the same 64 applicable cases. Study I varies evidence bundles, and Study II varies citation bindings and identifier spelling. Within each case and condition, both readers and the deterministic comparator receive the same packet. References derived from full logs are used only for scoring. Reference agreement and support from the visible evidence are assessed separately.
Level
Done
Del.
Inc.
Non.
Unr.
Elig.
A0
48
0
1
26
21
0
A1
216
189
10
123
83
2
A2
216
189
8
130
78
1
A3
214
0
4
115
95
0
Total
694
378
23
394
277
3
Table 1: Main campaign outcomes. Assigned: 696; completed: 694. Del. denotes actual delivery; Inc., Non., and Unr. denote fixed-rule incident, nonincident, and unresolved; Elig. denotes source eligibility.
Field or structure
Meaning and boundary
case_id
Packet identity; manifest links to the saved run.
target.action_id
Selected action; ordered tool log defines the earlier-record boundary.
stated_refs
Saved Basis: null means absent; an empty array means recorded empty.
id_mapping
Rows contain only id and locator ; present in C10/C11.
Record locator
Whole tool result, one-based list element, or request.user ; no substring offset.
native_inputs
R01/R11 provide input text, ID, locator, and channel. A user locator alone does not reveal user text.
Table 2: Implemented record-binding contract. IDs and locators are scoped to one saved execution; the case manifest identifies that execution. A mapping row supplies a relation, not record text.
Condition
D
A
+
∅
M
R00, R01
0
64
0
0
0
R10
36
28
0
18
18
R11
64
0
24
22
18
C00, C01
36
28
0
18
18
C10, C11
64
0
24
22
18
Table 3: Packet-only comparator, Q3, N=64 per condition. Every row has G=64 and U=0 . D includes correct nonempty sets ( + ), empty sets ( ∅ ), and explicit Basis-absent findings ( M ); A is justified abstention. Grouped conditions have identical counts.
Reader
Cond.
Q1 G
Q2
Q3 G
Q3 D
Q3 U
Luna
R00
63
60
64
0
0
Luna
R10
55
55
32
31
27
Luna
R01
63
64
64
0
0
Luna
R11
62
63
54
54
9
Sonnet
R00
64
64
64
0
0
Sonnet
R10
64
63
24
21
28
Table 4: Study I: all counts out of 64. Q1/Q3 G denotes grounded correctness; Q2 is exact complete-set accuracy. Q3 D denotes supported definite correctness and U unsupported assertions. R00: tools; R10: +actor account; R01: +native inputs; R11: both.
Reader
Cond.
N
A
D
U
W ∩ U
Luna
C00
28
7
0
21
12
Luna
C10
28
0
21
5
0
Luna
C01
28
14
0
14
11
Luna
C11
28
0
19
5
0
Sonnet
C00
28
2
0
26
22
Sonnet
C10
28
0
24
4
0
Table 5: Q3 in the 28 cases requiring a mapping. A is justified abstention; D is supported definite correctness; U is unsupported assertion; W ∩ U counts unsupported answers matching the complete reference. All 28 assignments per cell remain included.
Reader
Cond.
Q1 G
Q2
Q3 G
Q3 D
Q3 U
Q3 W
Luna
C00
55
49
38
31
21
43
Luna
C10
54
52
49
49
9
49
Luna
C01
60
48
42
28
15
39
Luna
C11
54
50
48
48
6
48
Sonnet
C00
64
64
26
24
28
46
Sonnet
C10
63
62
49
49
9
49
Table 6: Study II: all counts out of 64. G, D, and U follow Table 4 ; W is complete-reference agreement, which can include unsupported guesses. C00/C10: original IDs; C01/C11: opaque IDs; C10/C11 add the matching table. Q1 requires abstention throughout.
Study / endpoint
Reader
Contrast
Change (pp)
Interval (pp)
I: Q2
Luna
R10 − R00
-7.18
[-20.51, +3.85]
I: Q2
Luna
R11 − R01
-1.28
[-3.85, +0.00]
I: Q2
Sonnet
R10 − R00
-1.54
[-4.62, +0.00]
I: Q2
Sonnet
R11 − R01
+0.00
[+0.00, +0.00]
II: Q3
Luna
C10 − C00
+14.36
[+0.26, +30.00]
II: Q3
Luna
C11 − C01
+13.08
[-4.62, +32.56]
Table 7: Primary paired contrasts for each study, in percentage points. Tasks receive equal weight (13 clusters); 95% percentile intervals use 10,000 task-bootstrap draws. These exploratory intervals do not establish general harm, equivalence, or confirmatory significance.
Figure 2: Study II Q3 outcomes, retaining all 64 assignments per condition and reader. Grounded correctness is the sum of supported definite correctness and justified abstention, not a synonym for resolved sources. C00/C10 retain original IDs; C01/C11 use opaque IDs. C10 and C11 supply the corresponding complete binding table. The same-packet deterministic comparator is grounded-correct throughout.
Question
Supported finding
Q1: Request
Unresolved in both packets: authentic request text is withheld.
Q2: Occurrence
The recipient appears in act-01.result before the transfer in both packets.
Q3: Citation
Unresolved in C00; the explicit binding in C10 identifies the bill result as the saved Basis reference.
Table 8: Findings supported by the C00/C10 packets for rcase-001 . These summarize the existing illustrative case under the declared evidence contract.
Asynchronous monitoring, incident investigations, and compliance audits primarily rely on agent traces to reconstruct what happened. These analyses assume that LLM agents cannot tamper with their own execution traces. We show that local LLM agents such as Claude Code, Codex, Antigravity, Open Code and Grok Build fail to enforce this boundary. All tested harnesses, except Muse Code, allowed agents to delete their traces when asked, without triggering monitor guardrails. We also validate that external attackers can exploit this gap to induce trace deletion. Finally, we show that trace tampering behavior emerges naturally in frontier models, when agents try to improve their rewards. We advise practitioners to ensure trace logging happens through an independent interception mechanism outside of the agent's control, preserving trace integrity even in cases of full host compromise. Overall, our findings identify a concrete failure of trace integrity in agent infrastructure which can be used to conceal misaligned behaviors like scheming or sabotage.
Jeremy Qin, David Schmotz, Derck Prinzhorn +3
ELLIS Institute Tübingen · Max Planck Institute for Intelligent Systems · Tübingen AI Center +3
When an agent reuses logged experience, a changed decision may reflect the recorded score, the action's name, or the record's position in storage. Standard memory evaluations do not reveal which relationship drives that change. We introduce ReplayLens, a black-box audit that changes one relationship in the stored history at a time, holds the remaining interface fixed, and measures the resulting decision. Four interventions target four relationships. Outcome reassignment swaps which scores belong to which actions. Pair transport moves intact action-score pairs to new record slots. Consistent renaming relabels actions in both history and menu. Key-slot reassignment changes both score attachment and position. A constructive separation shows why the audit is needed: two memory writers with identical endpoint accuracy respond differently to the same replay, so conventional evaluation cannot resolve the underlying dependence. On black-box LLM interfaces, swapping scores changes decisions while moving intact pairs does not, separating score attachment from record order. A bounded-memory study exposes ingestion-order sensitivity that endpoint comparison misses. In sequential experiment planning, altered historical scores redirect exploration and reduce final utility despite fresh measurements. A code-debugging agent with sealed hidden tests shows the same pattern outside model selection. ReplayLens provides a relationship-level audit for deciding whether logged experience can be merged, reordered, or reindexed safely.
Dong Xu, Zhangfan Yang, Jiantao Wu +5
School of Artificial Intelligence, Shenzhen University · EasternDawn · School of Computer Science, University of Nottingham Ningbo
LLMs increasingly operate through coding-agent harnesses that inspect repositories, invoke tools, and modify files. Substituting the model behind such an agent can therefore change security-relevant decisions, including whether it verifies changes or recovers safely from failures. Existing LLM fingerprints largely infer identity from direct text or token distributions. In coding agents, these signals are mediated by system instructions, controller logic, tools, and execution feedback, limiting their transfer. We present LIDAR (LLM Identification from Decisions and Actions at Runtime), an active black-box fingerprinting method for coding-agent execution. Three coding probe pairs expose post-edit verification, transient-failure recovery, and specification--test conflict resolution under controlled changes. LIDAR represents the resulting trajectories with complementary instance-level and distribution-level features and compares them with clean references using a lightweight probabilistic identifier. It requires no access to model weights, logits, or provider internals. Across 36 models from seven families and two agent harnesses, LIDAR achieves high Top-1 accuracy and MRR and outperforms four existing fingerprinting and API-auditing baselines. Ablations confirm that the two feature levels, all probe pairs, and their controlled variants contribute. These results show that agent execution behavior provides model-identity evidence beyond final outputs.
Chuyi Wang, Xiaohui Xie, Tongze Wang +2
Department of Computer Science and Technology, Tsinghua University