When Retrieval Metrics Mislead: Measuring Policy Signal in Long-Horizon Tool-Use Agents
Organizations: Amazon Web Services
Abstract
Exact-match retrieval recall is often used as a proxy for whether a retriever supplies useful policy context to a downstream decision model. We test this proxy for pre-action policy classification in -bench using Qwen2.5-3B/7B classifiers. Under gold-policy conditioning, a compact structured state improves macro-F1 over raw trajectories by after tuning at 3B, with the same ordering at 7B under shared hyperparameters. We then replace the benchmark-designated governing rule with the top-ranked benchmark assertion retrieved from decision-time context. Although the exact governing rule is retrieved at rank 1 for only of airline states, the primary 3B classifier obtains macro-F1 with retrieved assertions versus with the gold rule (, task-cluster 95% CI ); random non-gold and no-assertion controls score and . We do not detect a macro-F1 difference between retrieved assertions and the gold rule in this configuration, although the interval remains too wide to establish non-inferiority. The same qualitative pattern appears with a second retriever and at 7B, while varying across fine-tuning configurations. These results show that exact-match recovery of the benchmark-designated rule can underestimate the downstream utility of retrieved benchmark assertions in this setting. Retrieval should therefore be evaluated inside the classification loop rather than by exact-match recall alone.