PathAnchor: Path-Structured Evidence for Scientific Agents
Authors: Qiuhui Chen, Jiafan Lu, Shuaimin Tang, Tao Dai, Suyuan Wang, Chenrui Ji, Zhenglei Zhou, Weimin Zhong
Organizations: School of Information Science and Engineering, East China University of Science and Technology · College of Intelligent Robotics and Advanced Manufacturing, Fudan University · Tencent
Scientific agents can retrieve relevant passages yet still lose functional order, mix evidence across sources, or state conclusions that exceed the retrieved record. We introduce PathAnchor, a bounded scientific reasoning system built on path-structured evidence workspaces. Instead of treating passages or extracted concepts as independent units, the system retrieves source-linked Material-Sensor-Signal-System trajectories that preserve role, direction, and the evidence supporting each transition. A controller uses three read-only tools to search paper-specific trajectories, trace paths across candidate sources, and open exact evidence before producing a claim-cited answer and an explicit evidence boundary. On 120 single- and cross-paper flexible-sensor questions, PathAnchor scores 82.6% and leads six evaluated systems. Under a matched controller, corpus, and six-call budget, replacing unordered concept graphs with path-structured records raises source recall from 61.3% to 82.9%, increases answers whose claims all cite opened evidence from 69.2% to 90.0%, and reduces tool calls. These results show that evidence organization affects retrieval and citation completeness under fixed agent resources.
Figures & tables
Figure 1: Illustrative comparison. P1–P4 denote distinct source papers. Flat retrieval can create an unsupported cross-paper link; PathAnchor preserves source-linked functional order, returns a cited supported claim, and explicitly marks evidence not found in the opened record.
Figure 2: PathAnchor architecture for FS EvidenceQA. A single controller navigates paper-specific MS 3 paths, opens linked article records, and returns claim-cited output with an explicit evidence boundary. E1/E2 are schematic; only opened article records count as final citations.
System prompt — PathAnchor
Tools and budget. Begin without candidate papers. Use search_ms3 , trace_paths , and inspect_evidence with at most six calls. Evidence rule. MS 3 path and mechanism IDs route retrieval but are not final proof. Every paper-specific claim cites an opened paper_id:E# article record; mark requests lacking support as insufficient.
Single-paper case — nv3_single_01_01
Question. For a SWCNT–RGO/PU fibre, how do TPU sheathing and 50 wt% RGO affect sensing, and what was not validated? Actions. Search → trace (P_081737)→ inspect (E1–E6) . Answer. After 60 min of ultrasonication, ΔR/R0 rose by 10 with TPU versus 690 without [E4]. For 50 wt% RGO, gauge factors were 8.76, 43.8, and 114.7 over 1–150%, 150–300%, and 300–465% strain [E3], with more than 1000 cycles [E6]. Opened records do not establish performance beyond 465% strain or for other RGO compositions. PARTIALLY_SUPPORTED .
Cross-paper case — nv3_cross_01_08
Question. Can a protective mouthguard also report deformation using a ZnO/GNP strain film? Actions. Search twice → trace four candidate papers → inspect eight article records. Answer. A mouthguard reduced transmitted incisor strain by over 55% [ P_095797:E5 ]; a separate ZnO/GNP film measured strain over a 44% range [ P_026428:E5 ]. Their paper-specific paths keep these results distinct: integration is a proposed design, not a tested device. PARTIALLY_SUPPORTED .
Table 1: Abridged PathAnchor prompt and two logged cases. The single-paper case uses P_081737 [ 19 ] ; the cross-paper case uses P_095797 and P_026428 [ 20 , 21 ] . Each paper_id:E# denotes an opened article record.
Metric
DSeek
Kimi
GLM
PQA2
OSch.
Ours
FS All
66.4
77.2
72.0
67.9
55.2
82.6
FS Single
78.8
87.6
84.4
76.0
60.0
91.9
FS Cross
54.1
66.9
59.7
59.8
50.4
73.3
FS Unsup. ↓
1.917
1.525
1.683
2.217
2.908
1.267
LQA
32.0
33.3
16.0
4.0
17.3
44.0
SQA
31.9
49.5
43.5
31.1
54.0
66.0
Table 2: Overall scores (%); FS Unsup. is the judge-rated count per answer. Ours denotes task-adapted MS 3 workflows on the three public tasks and PathAnchor on FS EvidenceQA. DSeek, PQA2, and OSch. are DeepSeek V4 Pro + Search, PaperQA2, and OpenScholar.
Figure 3: Paired outcomes and matched representation control. (a) FS EvidenceQA: PathAnchor (PA) versus each baseline; † marks a paired 95% interval crossing zero. (b) Source recall versus workspace size, with the proportion of answers whose claims all cite opened evidence. (c) Mean calls by action; B=6 is the budget-hit rate.
Question stratum
n
Text
MS 3
Δ [95% CI]
Covered, all
120
.688
.754
+.066 [.021, .112]
Single-paper
60
.772
.814
+.042 [-.022, .108]
Cross-paper
60
.604
.693
+.089 [.027, .152]
Partially covered
30
.765
.740
-.024 [-.093, .039]
Out-of-domain
30
.559
.587
+.028 [-.035, .092]
Table 3: Retrieval nDCG@10 (method-blind model relevance). Δ is MS 3 minus text FTS with paired 95% CI.
Department of Chemistry and Materials Science, Xi’an Jiaotong-Liverpool University, Suzhou 215123, Jiangsu, P. R. China · Suzhou Lab, Suzhou, P. R. China · College of Carbon Neutrality Future Technology, State Key Laboratory of Heavy Oil Processing, China University of Petroleum (Beijing), Beijing 102249, China +1