Auditing Agent Actions through Query-Conditioned Attribution
Organizations: University of Illinois Urbana-Champaign · IBM
Abstract
LLM agents increasingly take consequential actions through interactions with users, policies, and external tools. Auditing these agents requires automated attribution of realized actions to their historical basis. However, existing attribution formulations do not provide question-specific traces for diverse auditing objectives. Additionally, when access to the acting model is limited (e.g., in API-only deployments), applicable methods commonly rely on costly input perturbations or external LLM analysis of complete trajectories. We therefore formulate query-conditioned agent action attribution, a new task that takes a natural-language auditing query as input and recovers the source and ordered intermediate evidence for the query-specified aspect of an action. We instantiate this task with , a benchmark comprising 1,396 auditing queries across policy basis, parameter provenance, failure propagation, and unsafe-behavior tracing. To enable efficient, query-specific attribution, we use small open-weight models as attribution proposers that combine query-conditioned gradient saliency with query-semantic relevance to rank history units. Our proposer consistently achieves stronger source and evidence rankings at lower inference cost than open-weight baselines, improving source MRR by up to 40.9% and evidence MAP by 42.1% with only two forward passes and one backward pass. Controlled evaluations confirm that our proposer improves attribution specificity by adapting its rankings to fine-grained changes in the auditing query. Building on a proposer ensemble, our end-to-end system surpasses the strongest frontier-model baseline in source accuracy (64.5% vs.\ 60.4%) while reducing empirical deployment latency by 29.9% relative to the fastest frontier API baseline. Code and data will be released after the initial review period following final validation and cleanup.
Figures & tables
| Configuration | Qwen3.5-2B | Granite 3.3-2B | SmolLM3-3B | |||
| MRR | MAP | MRR | MAP | MRR | MAP | |
| Gradient sum | 0.4088 | 0.5702 | 0.4767 | 0.6080 | 0.4478 | 0.5818 |
| + norm. | 0.3701_{{\color[rgb]{0.6992,0.2266,0.2813}\scriptscriptstyle-.0387}} | 0.5908_{{\color[rgb]{0.0938,0.4727,0.3047}\scriptscriptstyle+.0206}} | 0.5511_{{\color[rgb]{0.0938,0.4727,0.3047}\scriptscriptstyle+.0744}} | 0.7208_{{\color[rgb]{0.0938,0.4727,0.3047}\scriptscriptstyle+.1128}} | 0.5720_{{\color[rgb]{0.0938,0.4727,0.3047}\scriptscriptstyle+.1242}} | 0.7132_{{\color[rgb]{0.0938,0.4727,0.3047}\scriptscriptstyle+.1314}} |
| +Query Grad. | 0.5007_{{\color[rgb]{0.0938,0.4727,0.3047}\scriptscriptstyle+.1306}} | \mathbf{0.6864}_{{\color[rgb]{0.0938,0.4727,0.3047}\scriptscriptstyle+.0956}} | 0.6595_{{\color[rgb]{0.0938,0.4727,0.3047}\scriptscriptstyle+.1084}} | \mathbf{0.7812}_{{\color[rgb]{0.0938,0.4727,0.3047}\scriptscriptstyle+.0604}} | 0.6628_{{\color[rgb]{0.0938,0.4727,0.3047}\scriptscriptstyle+.0908}} | \mathbf{0.7364}_{{\color[rgb]{0.0938,0.4727,0.3047}\scriptscriptstyle+.0232}} |
| +Query Rel. | \mathbf{0.5710}_{{\color[rgb]{0.0938,0.4727,0.3047}\scriptscriptstyle+.0703}} | 0.6848_{{\color[rgb]{0.6992,0.2266,0.2813}\scriptscriptstyle-.0016}} | \mathbf{0.6938}_{{\color[rgb]{0.0938,0.4727,0.3047}\scriptscriptstyle+.0343}} | 0.7388_{{\color[rgb]{0.6992,0.2266,0.2813}\scriptscriptstyle-.0425}} | \mathbf{0.6677}_{{\color[rgb]{0.0938,0.4727,0.3047}\scriptscriptstyle+.0049}} | 0.7247_{{\color[rgb]{0.6992,0.2266,0.2813}\scriptscriptstyle-.0117}} |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Acting Model | GPT-4.1 | Qwen3.5-9B |
| Queries | 1,141 | 255 |
| Trajectories | 871 | 214 |
| Task families | 335 | 140 |
| Domains | 7 | 7 |
| Policy basis | 290 | 64 |
| Parameter provenance | 290 | 65 |
| Source Hit@1 | Evidence Macro F1 | |||||
| Relation | Ours | Sonnet | Ours | Sonnet | ||
| Failure propagation | 0.416 | 0.222 | +0.195 | 0.673 | 0.576 | +0.097 |
| Parameter provenance | 0.365 | 0.677 | 0.312 | 0.530 | 0.615 | 0.085 |
| Policy basis | 0.859 | 0.543 | +0.316 | 0.653 | 0.620 | +0.033 |
| Unsafe-behavior tracing | 0.948 | 0.988 | 0.040 | 0.854 | 0.948 | 0.094 |
| Relation | Source errors (%) | Evidence F1 (%) under correct source | |||
| Rate | Proposer | Resolver | |||
| Failure propagation | 257 | 58.4 | 44.0 | 56.0 | 98.7 |
| Parameter provenance | 266 | 63.5 | 87.0 | 13.0 | 80.1 |
| Policy basis | 269 | 14.1 | 10.5 | 89.5 | 64.6 |
| Unsafe-behavior tracing | 249 | 5.2 | 7.7 | 92.3 | 87.4 |
| Overall | 1041 | 35.5 | 58.9 | 41.1 | 80.3 |
| Included proposer | Source selection | Evidence ranking | |||||
| Qwen-2B | Granite-2B | SmolLM3-3B | Agreement coverage | Accuracy under agreement | Disagreement Top-3 | R@ | MAP |
| ✓ | 100.00 | 37.75 | – | 0.5659 | 0.6848 | ||
| ✓ | ✓ | 50.05 | 70.06 | 80.58 | 0.5970 | 0.7110 | |
| ✓ | ✓ | ✓ | 37.85 | 77.16 | 80.22 | 0.6554 | 0.7568 |
| Resolver / Rule | Overall | Disagree | Evid. F1 | Exact |
| Qwen3.5-2B | 0.446 | 0.247 | 0.622 | 0.215 |
| Qwen3.5-4B | 0.645 | 0.567 | 0.675 | 0.298 |
| Qwen3.5-9B | 0.624 | 0.535 | 0.641 | 0.282 |
| Borda fallback | 0.572 | 0.450 | 0.670 | 0.305 |
| Proposer | Source ranking | Intermediate evidence | Inference cost | ||||||
| Hit@1 | Hit@3 | Hit@5 | MRR | Recall@ | MAP | Mean Fwd. | Mean Bwd. | Wall s/item | |
| Qwen3.5-2B | |||||||||
| ContextCite | 0.4338 | 0.5840 | 0.6496 | 0.5459 | 0.4890 | 0.5499 | 64.00 | 0 | 9.82 |
| Leave-one-out | 0.4863 | 0.6627 | 0.7199 | 0.5980 | 0.4591 | 0.5260 | 44.30 | 0 | 10.59 |
| PrefixDelta | 0.2884 | 0.6293 | 0.6639 | 0.4755 | 0.3188 | 0.4291 | 44.30 | 0 | 6.79 |
| Ours | 0.4136 | 0.7449 | 0.8772 | 0.6037 | 0.5609 | 0.6928 | 2.00 | 1 | 2.41 |
| Method | Source attribution | Intermediate evidence recovery | |||
| Hit@1 | Macro Hit@1 | Macro F1 | Micro F1 | Exact | |
| Gemini 2.5 Pro | [0.524, 0.588] | [0.533, 0.586] | [0.620, 0.661] | [0.579, 0.631] | [0.247, 0.307] |
| Claude Sonnet 4.6 | [0.573, 0.636] | [0.582, 0.632] | [0.669, 0.703] | [0.599, 0.647] | [0.271, 0.331] |
| Claude Haiku 4.5 | [0.374, 0.437] | [0.382, 0.435] | [0.429, 0.482] | [0.416, 0.481] | [0.187, 0.242] |
| GPT-4.1 | [0.550, 0.617] | [0.558, 0.616] | [0.535, 0.583] | [0.480, 0.536] | [0.247, 0.306] |
| Ensemble ( ) | [0.612, 0.678] | [0.621, 0.674] | [0.653, 0.696] | [0.572, 0.626] | [0.267, 0.329] |