physics.acc-phJul 27, 2026

A corrective agentic hybrid RAG and an operations-grounded evaluation for a scientific facility

Authors: Rajat SainjuDariusz JaroszHairong ShangMichael PrinceRyan M. AydelottMathew J. CherukaraYine SunMichael D. Borland

Organizations: Advanced Photon Source, Argonne National Laboratory, Lemont, Illinois 60439, USA

Abstract

Scientific user facilities accumulate decades of operational knowledge that no single search index covers: electronic logbooks, technical documents, internal wikis, operations chat messages, maintenance records, and live control-system data. We present APS-RAG, Advanced Photon Source Retrieval Augmented Generation, a deployed platform that makes the institutional knowledge at the Advanced Photon Source (APS) accessible to staff through natural-language queries, along with an operations-grounded evaluation. The retrieval engine fuses dense, sparse, and knowledge-graph (KG) channels with query-type-adaptive reciprocal-rank fusion, adds a corrective agentic loop, and runs a native-tool ReAct executor over a Model Context Protocol (MCP) tooling layer. We construct APS-Bench, a 50-question, question-answering (QA) dataset with auditable gold answers. Every retrieval-augmented variant numerically improves strict vital-nugget recall over a naive BM25 baseline (63.8%), with the full corrective Agentic GraphRAG scoring (70.3%). The cross-encoder reranker contributes significantly to answer quality: removing it and allowing the LLM to score relevance drastically reduces strict vital recall by 32.8%. The graph channel and corrective loop contribute positively as expected, but the performance gains are marginal. Additionally, we also compare the performance of open-source and closed-source LLMs in final answer synthesis. We release the APS-Bench construction methodology, the six-layer evaluation harness, and the underlying codebase, along with the '/aps-rag' retrieval agent skill framework, to support reproduction and adoption at other facilities. Together, the deployed platform and its operations-grounded evaluation present a promising workflow for trustworthy, statistically grounded AI assistance in facility operations, transferable to other large scientific instruments.

Explore similar work

CardsList
  1. SciRet: A Compute-Aware Empirical Study of Retrieval and Reranking for Scientific RAG

    Aug 4, 2026Kaysarul Anas Apurba, Md. Hasibul Hasan, Rofiqul Alam Shehab +1RerankingScientific Discovery