cs.CROct 7, 2026

Correct Answers, Unsupported Findings: Evidence Binding in Forensic Reconstruction of LLM Agent Logs

Authors: Taehyeon Yun, Dongho Kim, Geonwoo Kim, Juyoung Seo, Minseok Hur, Moohong Min

Organizations: aSungkyunkwan University, Republic of Korea

Abstract

Forensic reconstruction of LLM-agent actions requires not only recovering the correct value, but establishing which preserved record supports that finding. Tool logs, generated explanations, and local citation identifiers capture different parts of this evidence, yet a citation identifier does not establish a source unless its binding to a record is preserved. We audit this distinction using 64 mechanically checkable cases from saved AgentDojo Banking executions. Two LLM readers reconstruct source relationships under controlled variations in visible evidence and identifier-to-record bindings. We separately evaluate complete-record agreement, evidence-grounded findings, justified abstention, and unsupported assertions. With original identifiers and no binding table, Sonnet recovered every literal source location but made unsupported citation-source assertions in 26 of 28 cases requiring the missing relation; 22 nevertheless matched the complete reference. Adding explicit bindings improved grounded reconstruction for both readers, whereas identifier renaming alone provided no consistent remedy. A deterministic same-packet comparator correctly resolved the bounded task or abstained throughout. These results show that factual agreement alone is insufficient for evaluating forensic reconstruction of agent logs and motivate preserving explicit record bindings to distinguish supported findings from correct guesses.

Figures & tables

Explore similar work

CardsList
  1. LLM Agents Can Easily Tamper With Their Own Traces

    Sep 24, 2026Jeremy Qin, David Schmotz, Derck Prinzhorn +3TracesLarge Language Model Agents

  2. ReplayLens: Auditing Agents' Use of Outcomes

    Sep 28, 2026Dong Xu, Zhangfan Yang, Jiantao Wu +5ReplayInteraction History

  3. Who Is Behind the Harness? Fingerprinting LLMs through Agentic Behavior

    Sep 23, 2026Chuyi Wang, Xiaohui Xie, Tongze Wang +2FingerprintCoding Agents