cs.CLJun 22, 2026

When Retrieval Metrics Mislead: Measuring Policy Signal in Long-Horizon Tool-Use Agents

Authors: Tianyu DingJuan Pablo De la Cruz Weinstein

Organizations: Amazon Web Services

Abstract

Exact-match retrieval recall is often used as a proxy for whether a retriever supplies useful policy context to a downstream decision model. We test this proxy for pre-action policy classification in ττ-bench using Qwen2.5-3B/7B classifiers. Under gold-policy conditioning, a compact structured state improves macro-F1 over raw trajectories by 0.200.20 after tuning at 3B, with the same ordering at 7B under shared hyperparameters. We then replace the benchmark-designated governing rule with the top-ranked benchmark assertion retrieved from decision-time context. Although the exact governing rule is retrieved at rank 1 for only 7%7\% of airline states, the primary 3B classifier obtains macro-F1 0.580.58 with retrieved assertions versus 0.600.60 with the gold rule (Δ=0.02Δ=-0.02, task-cluster 95% CI [0.23,+0.21][-0.23,+0.21]); random non-gold and no-assertion controls score 0.320.32 and 0.210.21. We do not detect a macro-F1 difference between retrieved assertions and the gold rule in this configuration, although the interval remains too wide to establish non-inferiority. The same qualitative pattern appears with a second retriever and at 7B, while varying across fine-tuning configurations. These results show that exact-match recovery of the benchmark-designated rule can underestimate the downstream utility of retrieved benchmark assertions in this setting. Retrieval should therefore be evaluated inside the classification loop rather than by exact-match recall alone.

Explore similar work

CardsList