Forensic reconstruction of LLM-agent actions requires not only recovering the correct value, but establishing which preserved record supports that finding. Tool logs, generated explanations, and local citation identifiers capture different parts of this evidence, yet a citation identifier does not establish a source unless its binding to a record is preserved. We audit this distinction using 64 mechanically checkable cases from saved AgentDojo Banking executions. Two LLM readers reconstruct source relationships under controlled variations in visible evidence and identifier-to-record bindings. We separately evaluate complete-record agreement, evidence-grounded findings, justified abstention, and unsupported assertions. With original identifiers and no binding table, Sonnet recovered every literal source location but made unsupported citation-source assertions in 26 of 28 cases requiring the missing relation; 22 nevertheless matched the complete reference. Adding explicit bindings improved grounded reconstruction for both readers, whereas identifier renaming alone provided no consistent remedy. A deterministic same-packet comparator correctly resolved the bounded task or abstained throughout. These results show that factual agreement alone is insufficient for evaluating forensic reconstruction of agent logs and motivate preserving explicit record bindings to distinguish supported findings from correct guesses.
Figures & tables
Figure 1: Both studies reuse the same 64 applicable cases. Study I varies evidence bundles, and Study II varies citation bindings and identifier spelling. Within each case and condition, both readers and the deterministic comparator receive the same packet. References derived from full logs are used only for scoring. Reference agreement and support from the visible evidence are assessed separately.
Level
Done
Del.
Inc.
Non.
Unr.
Elig.
A0
48
0
1
26
21
0
A1
216
189
10
123
83
2
A2
216
189
8
130
78
1
A3
214
0
4
115
95
0
Total
694
378
23
394
277
3
Table 1: Main campaign outcomes. Assigned: 696; completed: 694. Del. denotes actual delivery; Inc., Non., and Unr. denote fixed-rule incident, nonincident, and unresolved; Elig. denotes source eligibility.
Field or structure
Meaning and boundary
case_id
Packet identity; manifest links to the saved run.
target.action_id
Selected action; ordered tool log defines the earlier-record boundary.
stated_refs
Saved Basis: null means absent; an empty array means recorded empty.
id_mapping
Rows contain only id and locator ; present in C10/C11.
Record locator
Whole tool result, one-based list element, or request.user ; no substring offset.
native_inputs
R01/R11 provide input text, ID, locator, and channel. A user locator alone does not reveal user text.
Table 2: Implemented record-binding contract. IDs and locators are scoped to one saved execution; the case manifest identifies that execution. A mapping row supplies a relation, not record text.
Condition
D
A
+
∅
M
R00, R01
0
64
0
0
0
R10
36
28
0
18
18
R11
64
0
24
22
18
C00, C01
36
28
0
18
18
C10, C11
64
0
24
22
18
Table 3: Packet-only comparator, Q3, N=64 per condition. Every row has G=64 and U=0 . D includes correct nonempty sets ( + ), empty sets ( ∅ ), and explicit Basis-absent findings ( M ); A is justified abstention. Grouped conditions have identical counts.
Reader
Cond.
Q1 G
Q2
Q3 G
Q3 D
Q3 U
Luna
R00
63
60
64
0
0
Luna
R10
55
55
32
31
27
Luna
R01
63
64
64
0
0
Luna
R11
62
63
54
54
9
Sonnet
R00
64
64
64
0
0
Sonnet
R10
64
63
24
21
28
Table 4: Study I: all counts out of 64. Q1/Q3 G denotes grounded correctness; Q2 is exact complete-set accuracy. Q3 D denotes supported definite correctness and U unsupported assertions. R00: tools; R10: +actor account; R01: +native inputs; R11: both.
Reader
Cond.
N
A
D
U
W ∩ U
Luna
C00
28
7
0
21
12
Luna
C10
28
0
21
5
0
Luna
C01
28
14
0
14
11
Luna
C11
28
0
19
5
0
Sonnet
C00
28
2
0
26
22
Sonnet
C10
28
0
24
4
0
Table 5: Q3 in the 28 cases requiring a mapping. A is justified abstention; D is supported definite correctness; U is unsupported assertion; W ∩ U counts unsupported answers matching the complete reference. All 28 assignments per cell remain included.
Reader
Cond.
Q1 G
Q2
Q3 G
Q3 D
Q3 U
Q3 W
Luna
C00
55
49
38
31
21
43
Luna
C10
54
52
49
49
9
49
Luna
C01
60
48
42
28
15
39
Luna
C11
54
50
48
48
6
48
Sonnet
C00
64
64
26
24
28
46
Sonnet
C10
63
62
49
49
9
49
Table 6: Study II: all counts out of 64. G, D, and U follow Table 4 ; W is complete-reference agreement, which can include unsupported guesses. C00/C10: original IDs; C01/C11: opaque IDs; C10/C11 add the matching table. Q1 requires abstention throughout.
Study / endpoint
Reader
Contrast
Change (pp)
Interval (pp)
I: Q2
Luna
R10 − R00
-7.18
[-20.51, +3.85]
I: Q2
Luna
R11 − R01
-1.28
[-3.85, +0.00]
I: Q2
Sonnet
R10 − R00
-1.54
[-4.62, +0.00]
I: Q2
Sonnet
R11 − R01
+0.00
[+0.00, +0.00]
II: Q3
Luna
C10 − C00
+14.36
[+0.26, +30.00]
II: Q3
Luna
C11 − C01
+13.08
[-4.62, +32.56]
Table 7: Primary paired contrasts for each study, in percentage points. Tasks receive equal weight (13 clusters); 95% percentile intervals use 10,000 task-bootstrap draws. These exploratory intervals do not establish general harm, equivalence, or confirmatory significance.
Figure 2: Study II Q3 outcomes, retaining all 64 assignments per condition and reader. Grounded correctness is the sum of supported definite correctness and justified abstention, not a synonym for resolved sources. C00/C10 retain original IDs; C01/C11 use opaque IDs. C10 and C11 supply the corresponding complete binding table. The same-packet deterministic comparator is grounded-correct throughout.
Question
Supported finding
Q1: Request
Unresolved in both packets: authentic request text is withheld.
Q2: Occurrence
The recipient appears in act-01.result before the transfer in both packets.
Q3: Citation
Unresolved in C00; the explicit binding in C10 identifies the bill result as the saved Basis reference.
Table 8: Findings supported by the C00/C10 packets for rcase-001 . These summarize the existing illustrative case under the declared evidence contract.