Conventional reasoning protocols present a fixed, preselected task, so they cannot test whether an agent propagates relevant updates, preserves unaffected work, or reconstructs a historical task binding. We therefore study \emph{dynamic task routing}, in which an event stream revises task bindings and a system must select the document version valid at each query time before solving it. To study this problem, we repurpose six widely used benchmarks: MMLU, MMLU-Pro, MedMCQA, MATH, GPQA, and HumanEval into 31{,}119 dynamic episodes comprising 373{,}428 temporally categorized queries. This setting exposes a central trade-off: recomputing after every event wastes work, whereas unguarded reuse returns stale conclusions. We introduce the Revision-Aware Independent Agent Graph (RIAG), a bounded multi-agent policy that separates deterministic temporal resolution from task reasoning. RIAG caches solutions by immutable document identity, starts each fresh task with two unexposed attempts, and conditionally invokes audit and repair, using at most four calls per document version. On this collection, homogeneous RIAG achieves 54.24% joint routing-and-answer accuracy at 0.62 calls/query, compared with 32.22% at 18.00 calls/query for the strongest comparison method; heterogeneous RIAG reaches 49.78% at 0.63 calls/query.
Figures & tables
Figure 1: Conventional isolated question and response protocol versus dynamic task routing protocol. The dynamic protocol retains an arrival ordered event history, distinguishes arrival time from effective time, routes each query to a fixed-content task version, restores an underlying route after expiration, and revises historical resolution when a previously trusted event is retracted.
Figure 2: RIAG workflow. A deterministic resolver selects the version valid at the query’s world time; exact version–configuration keys enable reuse; and cache misses open a two-to-four-call exposure-aware graph.
Accuracy (%) ↑
Cost ↓
Joint accuracy (%) by query category ↑
Joint
Answer
Selection
Calls/q
Baseline
Changed
Propag.
Preserv.
Historical
Debate
32.22
39.42
62.78
18.00
53.06
50.94
34.71
24.78
26.48
Self-Consistency
26.62
34.33
61.40
6.00
45.70
37.14
27.56
22.50
21.29
Refine
27.84
35.34
60.85
19.00
48.49
42.94
29.96
20.75
23.25
ReConcile
28.92
35.37
63.51
18.00
49.42
43.83
31.08
21.98
24.22
Mixture-of-Agents
28.98
36.37
61.06
19.00
50.30
47.25
32.15
20.88
23.42
Table 1: Performance on the dynamic task routing variants of MMLU, MMLU-Pro, MedMCQA, MATH, GPQA, and HumanEval. Dataset-wise breakdowns are provided in Appendix D.5 .
Figure 3: Analysis of the transferability of the proposed deterministic resolver design to Graph-of-Agents on MATH (500 episodes and 6,000 queries per configuration). Lines pair vanilla and resolver-controlled runs with the same pooling and model-pool labels. (a) Joint accuracy, with changes in percentage points; (b) selection accuracy; (c) recorded calls/query.
Figure 4: Reasoning geometry and conditional intervention. Left: task-centered response UMAP for RIAG invocations, jointly fitted across homogeneous and heterogeneous roles. Right: useful diversity by solver–checker distance quartile and outcomes conditional on auditing. Rescue means that two wrong initial answers become a correct final answer; damage means that at least one correct initial answer becomes a wrong final answer.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Total calls ↓
Calls/query ↓
Relative to RIAG ↓
Debate
6,721,704
18.00
28.90 ×
Self-Consistency
2,240,568
6.00
9.63 ×
Refine
7,095,132
19.00
30.50 ×
ReConcile
6,721,704
18.00
28.90 ×
Mixture-of-Agents
7,095,132
19.00
30.50 ×
Self-MoA
2,613,996
7.00
11.24 ×
Appendix
Table 2: Model calls on the full collection. The last column divides each method’s calls by those of homogeneous RIAG.
Metric
Homogeneous
Heterogeneous
Model calls ↓
232,614
234,436
Calls/query ↓
0.62
0.63
Input tokens ↓
98,070,477
106,642,017
Output tokens ↓
27,697,101
27,486,066
Output tokens/query ↓
74.2
73.6
Parse-failed queries, original run ↓
20,739
4,453
Appendix
Table 3: Recorded cost of the two full-collection RIAG production runs (31,119 episodes and 373,428 scheduled queries each; 16 workers; parallel independent calls enabled). Homogeneous totals include the 34,796 recovery calls of the repaired overlay (17.4M input and 4.9M output tokens); its wall-clock covers the production phase only.
Multi-Domain ↑
Domain-Specific ↑
MMLU
MMLU-Pro
GPQA
MATH
HumanEval
MedMCQA
All ↑
Debate
38.66
25.76
23.02
29.60
14.53
30.66
32.22
Self-Consistency
34.28
18.62
21.51
2.93
25.41
27.02
26.62
Refine
34.58
20.73
18.73
20.93
22.26
27.14
27.84
ReConcile
35.44
21.61
18.77
27.60
26.88
28.80
28.92
Mixture-of-Agents
35.85
21.70
18.77
25.77
22.36
27.98
28.98
Appendix
Table 4: Joint accuracy by dataset on the full collection; each column uses every query of its dataset (GPQA 2,376; MATH 6,000; HumanEval 1,968; MedMCQA 50,196; MMLU-Pro 144,384; MMLU 168,504).
Multi-Domain ↑
Domain-Specific ↑
MMLU
MMLU-Pro
GPQA
MATH
HumanEval
MedMCQA
All ↑
Debate
48.41
29.60
31.99
30.05
14.53
39.92
39.42
Self-Consistency
44.79
22.80
30.56
3.17
25.46
36.60
34.33
Refine
44.81
24.70
26.73
21.33
22.31
36.70
35.34
ReConcile
44.25
24.89
26.30
28.00
26.88
37.33
35.37
Mixture-of-Agents
45.96
25.45
27.06
26.22
22.56
37.80
36.37
Appendix
Table 5: Answer accuracy by dataset on the full collection; each column uses every query of its dataset (GPQA 2,376; MATH 6,000; HumanEval 1,968; MedMCQA 50,196; MMLU-Pro 144,384; MMLU 168,504). The RIAG rows equal their joint rows because the RIAG artifacts record one exact-correctness predicate.
Multi-Domain ↑
Domain-Specific ↑
MMLU
MMLU-Pro
GPQA
MATH
HumanEval
MedMCQA
All ↑
Debate
62.08
63.28
62.75
63.17
51.27
64.13
62.78
Self-Consistency
61.07
61.67
59.43
58.62
54.27
62.47
61.40
Refine
60.70
60.88
59.26
59.32
45.53
62.09
60.85
ReConcile
62.57
64.32
62.79
64.30
64.43
64.29
63.51
Mixture-of-Agents
60.85
61.26
59.93
59.80
42.73
62.11
61.06
Appendix
Table 6: Selection accuracy by dataset on the full collection.
Figure 5: RIAG on the MATH split with one setting varied at a time (500 episodes, 6,000 queries per run); stars mark the reference. Rows: joint accuracy, model calls per query, and output tokens per query. Labels give the two accuracy changes whose paired 95% intervals exclude zero.
Figure 6: Reference-setting mechanism diagnostics. (a) Cache-hit rate by query category. (b) Joint accuracy by category with episode-cluster bootstrap intervals. (c) Route depth among unique fresh task executions. (d) Route-conditioned accuracy; numeric labels give task counts. Panels (c,d) describe the policy’s selection behavior and are not causal effects of additional calls.
Figure 7: Matched solver–checker diversity on 1,386 unique MATH tasks per configuration. (a) Original-space cosine-distance distributions; the annotation reports the paired heterogeneous-minus-homogeneous difference and a 10,000-resample task-bootstrap interval. (b) Initial-pair outcomes and audit frequency. “One correct” denotes useful diversity; “wrong agreement” denotes equal normalized answers that are both incorrect.
Figure 8: Initial-pair outcomes on a fixed display subset of the task-centered UMAP. Each segment links the solver and checker responses for one task, and the solver endpoint is colored by normalized answer outcome. All matched pairs enter the numerical analysis; the plotted subset is for readability and is not a learned decision boundary.
Stage
Operation
Control
Source freeze
Snapshot the six public benchmarks also evaluated by Graph-of-Agents.
Preserve source bytes, row order, split labels, counts, and hashes.
AI-assisted design
Use Astra to help develop and inspect the dynamic event-stream transformation.
Do not use model output to rewrite source tasks or establish gold answers.
Deterministic build
Map every source row to one episode; sample two within-dataset partners; generate 12 temporal queries.
Three reviewers inspect samples from all resulting dataset families.
Check source fidelity, temporal semantics, document–gold binding, and public/private separation.
Mechanical audit
Replay temporal selection with an independent validator and run regression tests.
Verify 31,119 episodes, 373,428 queries, schemas, manifests, and artifact checksums.
Appendix
Table 7: Dataset curation and quality-control stages. AI assistance supports protocol development, while deterministic generation and independent replay define the released artifacts and their gold bindings.
#
Public update before the query
a
w
Checkpoint
Selected
MMLU — mmlu-task-version-000000 . Documents: A = row 3370 (commander analogy); B = row 11280 (new-van liability); C = row 0 (field-extension degree).
1
Load B, A, and C; bind requests; activate C
2
2
Baseline
C
2
Supersede the active binding with B
7
7
Changed
B
3
Add an unrelated request; query the archive
12
12
Preservation
C
4
No new update; query the earlier world state
12
2
Historical
C
5
Schedule C for e=28
18
18
Preservation
B
Appendix
Table 8: Six actual episodes.
Dataset
Local split
Source items / episodes
Queries
Episode share
MMLU
test
14,042
168,504
45.1%
MMLU-Pro
test
12,032
144,384
38.7%
MedMCQA
test
4,183
50,196
13.4%
MATH
test
500
6,000
1.6%
GPQA
test
198
2,376
0.6%
HumanEval
dev
164
1,968
0.5%
Appendix
Table 9: The dynamic task routing collection. Every source item anchors one episode and every episode contains 12 checkpoint queries.
Dataset
Frozen export
Task field
Valid answer
Evaluator
MMLU
test/MMLU_test.json
question
A–D
Exact source option letter
MMLU-Pro
test/MMLU_Pro_test.json
question
A–J
Exact source option letter
MedMCQA
test/MedMCQA_test.json
question
A–D
Exact source option letter
GPQA
test/GPQA_test.json
question
A–D
Exact source option letter
MATH
test/MATH_test.json
question
String
Exact match to gold_answer
HumanEval
dev/human_eval_dev.json
prompt
Python
Hidden test with original entry_point
Appendix
Table 10: Dataset-specific input and output contracts. The public document is the indicated source field without rewriting. All final predictions also carry a document label in document_id .
Figure 9: Protocol-design comparison over the same source inventory. The top-left panel counts documents, checkpoint queries, event records, and public stream records per item or episode. Context lengths are medians of whitespace-separated visible tokens. The capability matrix states what is explicitly probed by the protocol, not measured model performance.
Figure 10: Additional diagnostic signal from Dynamic task routing. Left: one isolated response becomes 12 categorized checkpoints. Right: fraction of exact-answer episodes whose baseline and changed routes select different documents but share an identical gold answer string.
Figure 11: Semantic coverage of the 31,119 unique source tasks. (a) UMAP projection colored by source dataset. Axes and long-range projected distances have no direct semantic interpretation. (b) Dataset composition of each task’s 15 nearest neighbors, computed in the original normalized embedding space.
Figure 12: Mean pairwise cosine distance among the three source documents in each observed episode, compared with a within-dataset random-partner control. Distances are measured in the original normalized embedding space; boxes summarize episodes within each source dataset.
Figure 13: Descriptive semantic outcome landscapes for complete-run artifacts. (a) Homogeneous RIAG joint accuracy. (b) Heterogeneous minus homogeneous RIAG joint accuracy. (c) Homogeneous GoA mean minus max pooling. (d) Homogeneous GoA-mean joint accuracy. Colors are 50-neighbor averages in the original embedding space displayed on the common UMAP projection; they are not independent tests or causal effects.
Figure 14: Selection and joint accuracy by quartile of episode document separation, from closest (Q1) to farthest (Q4). Distance is the episode’s mean pairwise cosine distance in the original embedding space. Curves are descriptive aggregates over different mixtures of source datasets and do not identify a distance effect.
Q
Public update immediately before query
a
w
Request
Category
Gold pair
1
Load A, B, C; bind archive to C; set active task to C
6
6
Delivery
Baseline
(C,C)
2
Supersede the C assertion with active task B
8
8
Delivery
Changed
(B,D)
3
Add an unrelated request; query the fixed archive
10
10
Archive
Preservation
(C,C)
4
No update; query the delivery binding at world time 6
10
6
Delivery
Historical
(C,C)
5
Schedule active task A for effective time 21
14
14
Delivery
Preservation
(B,D)
6
No update; scheduled A is now effective
23
23
Delivery
Propagation
(A,A)
Appendix
Table 11: Complete context for the shared-episode example. Panel A gives all 12 queries in stream order; an update shown on a row arrives immediately before that query. Panel B gives the six recorded responses to highlighted query Q2 (ID c9e557d88cb3c4c275699e23 ). Correctness requires the complete (document, answer) pair, not a matching option symbol alone.
We introduce ARBIGRAPH, a benchmark generator for evaluating whether tool-assisted language agents can retain, update, compose, and discard task-relevant context across extended reasoning workflows. ARBIGRAPH represents each task as a natural-language problem with an executable Python solver, and composes tasks through typed intermediate states, instantiated here as scalar and list values. This design enables controllable task graphs whose length, dependency structure, distractor count, and value type can be varied while preserving exact automatic verification. We instantiate ARBIGRAPH with math, GSM-style word-problems, and Python-tracing task categories, and evaluate a Qwen3.5-27B tool-assisted agent across four topologies. The results show high accuracy on isolated tasks but substantial degradation on more complex dependent tasks: accuracy drops by up to 33.3% on branching chains of dependent math tasks. This shows that ARBIGRAPH exposes failures that are not visible from single-task evaluation alone. Our code, generated datasets, and evaluation results are available at https://github.com/pavelgolikov/ArbiGraph.git
Pavel Golikov, Evgenii Opryshko, Gennady Pekhimenko +1
Tackling complex reasoning tasks typically relies on massive monolithic LLMs, which suffer from severe computational redundancy. While task decomposition through structured pipelines or multi-agent collaborations offers an alternative, these approaches inevitably fall into a critical dilemma: predefined static topologies are highly vulnerable to cascading errors, whereas unconstrained dynamic agents suffer from trajectory divergence and unpredictable memory bloat. To address this, we present DynaGraph, a lightweight multi-model framework driven by dynamic topological reconfiguration. At the execution level, DynaGraph multiplexes time-division PEFT adapters over a shared base model, enabling both full system training and inference deployment on a single consumer-grade GPU. At the routing level, the Evaluator continuously monitors execution confidence to trigger hierarchical self-healing: Fine-grained Patching for localized data gaps and Subgraph Reconstruction for severe logical ruptures. Experiments on StrategyQA, MATH, and FinQA demonstrate our 8B model closely approximates the reasoning capabilities of a 72B monolithic model (e.g., 87.6% on StrategyQA, 82.7% on MATH). Furthermore, it reduces latency by up to 68.1% and token consumption by 68.6% compared to unconstrained dynamic architectures.
Yanxing Guo, Zihao Zheng, Fangzhou Wu +4
Peking University · Nanjing University · Beijing Advanced Innovation Center for Integrated Circuits +2
Agentic GraphRAG trains language-model agents to iteratively retrieve and reason over graph-structured evidence, enabling more accurate and context-aware decision-making by efficiently navigating complex information networks. However, outcome-only reinforcement learning suffers from \textit{\textbf{answer-path reward aliasing}}, where correct answers may come from shortcuts rather than useful evidence paths. It also exhibits \textit{\textbf{search-update ambiguity}}, as scalar trajectory-level feedback does not indicate which retrieval actions to adjust. To mitigate these shortcomings, we present PathRouter, a path-aware training framework for agentic GraphRAG. PathRouter jointly evaluates each trajectory along answer correctness and evidence-path overlap, yielding four trajectory categories with differentiated GRPO advantage scaling that suppresses shortcut reinforcement while preserving evidence-seeking behavior. For evidence-poor trajectories, a frozen gold-evidence teacher provides token-level KL guidance on reasoning and search-query tokens, excluding answer tokens to avoid direct response imitation. Experiments on six QA benchmarks across three model sizes show that PathRouter consistently improves answer F1 and evidence-path overlap, achieving average F1 gains of 3.1 on 3B and 4.9 on 7B models compared to a strong baseline.
Bo Wang, Heyan Huang, Yaolin Li +6
Beijing Institute of Technology · Joy Future Academy