Conventional reasoning protocols present a fixed, preselected task, so they cannot test whether an agent propagates relevant updates, preserves unaffected work, or reconstructs a historical task binding. We therefore study \emph{dynamic task routing}, in which an event stream revises task bindings and a system must select the document version valid at each query time before solving it. To study this problem, we repurpose six widely used benchmarks: MMLU, MMLU-Pro, MedMCQA, MATH, GPQA, and HumanEval into 31{,}119 dynamic episodes comprising 373{,}428 temporally categorized queries. This setting exposes a central trade-off: recomputing after every event wastes work, whereas unguarded reuse returns stale conclusions. We introduce the Revision-Aware Independent Agent Graph (RIAG), a bounded multi-agent policy that separates deterministic temporal resolution from task reasoning. RIAG caches solutions by immutable document identity, starts each fresh task with two unexposed attempts, and conditionally invokes audit and repair, using at most four calls per document version. On this collection, homogeneous RIAG achieves 54.24% joint routing-and-answer accuracy at 0.62 calls/query, compared with 32.22% at 18.00 calls/query for the strongest comparison method; heterogeneous RIAG reaches 49.78% at 0.63 calls/query.
Figures & tables
Figure 1: Conventional isolated question and response protocol versus dynamic task routing protocol. The dynamic protocol retains an arrival ordered event history, distinguishes arrival time from effective time, routes each query to a fixed-content task version, restores an underlying route after expiration, and revises historical resolution when a previously trusted event is retracted.
Figure 2: RIAG workflow. A deterministic resolver selects the version valid at the query’s world time; exact version–configuration keys enable reuse; and cache misses open a two-to-four-call exposure-aware graph.
Accuracy (%) ↑
Cost ↓
Joint accuracy (%) by query category ↑
Joint
Answer
Selection
Calls/q
Baseline
Changed
Propag.
Preserv.
Historical
Debate
32.22
39.42
62.78
18.00
53.06
50.94
34.71
24.78
26.48
Self-Consistency
26.62
34.33
61.40
6.00
45.70
37.14
27.56
22.50
21.29
Refine
27.84
35.34
60.85
19.00
48.49
42.94
29.96
20.75
23.25
ReConcile
28.92
35.37
63.51
18.00
49.42
43.83
31.08
21.98
24.22
Mixture-of-Agents
28.98
36.37
61.06
19.00
50.30
47.25
32.15
20.88
23.42
Table 1: Performance on the dynamic task routing variants of MMLU, MMLU-Pro, MedMCQA, MATH, GPQA, and HumanEval. Dataset-wise breakdowns are provided in Appendix D.5 .
Figure 3: Analysis of the transferability of the proposed deterministic resolver design to Graph-of-Agents on MATH (500 episodes and 6,000 queries per configuration). Lines pair vanilla and resolver-controlled runs with the same pooling and model-pool labels. (a) Joint accuracy, with changes in percentage points; (b) selection accuracy; (c) recorded calls/query.
Figure 4: Reasoning geometry and conditional intervention. Left: task-centered response UMAP for RIAG invocations, jointly fitted across homogeneous and heterogeneous roles. Right: useful diversity by solver–checker distance quartile and outcomes conditional on auditing. Rescue means that two wrong initial answers become a correct final answer; damage means that at least one correct initial answer becomes a wrong final answer.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Total calls ↓
Calls/query ↓
Relative to RIAG ↓
Debate
6,721,704
18.00
28.90 ×
Self-Consistency
2,240,568
6.00
9.63 ×
Refine
7,095,132
19.00
30.50 ×
ReConcile
6,721,704
18.00
28.90 ×
Mixture-of-Agents
7,095,132
19.00
30.50 ×
Self-MoA
2,613,996
7.00
11.24 ×
Appendix
Table 2: Model calls on the full collection. The last column divides each method’s calls by those of homogeneous RIAG.
Metric
Homogeneous
Heterogeneous
Model calls ↓
232,614
234,436
Calls/query ↓
0.62
0.63
Input tokens ↓
98,070,477
106,642,017
Output tokens ↓
27,697,101
27,486,066
Output tokens/query ↓
74.2
73.6
Parse-failed queries, original run ↓
20,739
4,453
Appendix
Table 3: Recorded cost of the two full-collection RIAG production runs (31,119 episodes and 373,428 scheduled queries each; 16 workers; parallel independent calls enabled). Homogeneous totals include the 34,796 recovery calls of the repaired overlay (17.4M input and 4.9M output tokens); its wall-clock covers the production phase only.
Multi-Domain ↑
Domain-Specific ↑
MMLU
MMLU-Pro
GPQA
MATH
HumanEval
MedMCQA
All ↑
Debate
38.66
25.76
23.02
29.60
14.53
30.66
32.22
Self-Consistency
34.28
18.62
21.51
2.93
25.41
27.02
26.62
Refine
34.58
20.73
18.73
20.93
22.26
27.14
27.84
ReConcile
35.44
21.61
18.77
27.60
26.88
28.80
28.92
Mixture-of-Agents
35.85
21.70
18.77
25.77
22.36
27.98
28.98
Appendix
Table 4: Joint accuracy by dataset on the full collection; each column uses every query of its dataset (GPQA 2,376; MATH 6,000; HumanEval 1,968; MedMCQA 50,196; MMLU-Pro 144,384; MMLU 168,504).
Multi-Domain ↑
Domain-Specific ↑
MMLU
MMLU-Pro
GPQA
MATH
HumanEval
MedMCQA
All ↑
Debate
48.41
29.60
31.99
30.05
14.53
39.92
39.42
Self-Consistency
44.79
22.80
30.56
3.17
25.46
36.60
34.33
Refine
44.81
24.70
26.73
21.33
22.31
36.70
35.34
ReConcile
44.25
24.89
26.30
28.00
26.88
37.33
35.37
Mixture-of-Agents
45.96
25.45
27.06
26.22
22.56
37.80
36.37
Appendix
Table 5: Answer accuracy by dataset on the full collection; each column uses every query of its dataset (GPQA 2,376; MATH 6,000; HumanEval 1,968; MedMCQA 50,196; MMLU-Pro 144,384; MMLU 168,504). The RIAG rows equal their joint rows because the RIAG artifacts record one exact-correctness predicate.
Multi-Domain ↑
Domain-Specific ↑
MMLU
MMLU-Pro
GPQA
MATH
HumanEval
MedMCQA
All ↑
Debate
62.08
63.28
62.75
63.17
51.27
64.13
62.78
Self-Consistency
61.07
61.67
59.43
58.62
54.27
62.47
61.40
Refine
60.70
60.88
59.26
59.32
45.53
62.09
60.85
ReConcile
62.57
64.32
62.79
64.30
64.43
64.29
63.51
Mixture-of-Agents
60.85
61.26
59.93
59.80
42.73
62.11
61.06
Appendix
Table 6: Selection accuracy by dataset on the full collection.
Figure 5: RIAG on the MATH split with one setting varied at a time (500 episodes, 6,000 queries per run); stars mark the reference. Rows: joint accuracy, model calls per query, and output tokens per query. Labels give the two accuracy changes whose paired 95% intervals exclude zero.
Figure 6: Reference-setting mechanism diagnostics. (a) Cache-hit rate by query category. (b) Joint accuracy by category with episode-cluster bootstrap intervals. (c) Route depth among unique fresh task executions. (d) Route-conditioned accuracy; numeric labels give task counts. Panels (c,d) describe the policy’s selection behavior and are not causal effects of additional calls.
Figure 7: Matched solver–checker diversity on 1,386 unique MATH tasks per configuration. (a) Original-space cosine-distance distributions; the annotation reports the paired heterogeneous-minus-homogeneous difference and a 10,000-resample task-bootstrap interval. (b) Initial-pair outcomes and audit frequency. “One correct” denotes useful diversity; “wrong agreement” denotes equal normalized answers that are both incorrect.
Figure 8: Initial-pair outcomes on a fixed display subset of the task-centered UMAP. Each segment links the solver and checker responses for one task, and the solver endpoint is colored by normalized answer outcome. All matched pairs enter the numerical analysis; the plotted subset is for readability and is not a learned decision boundary.
Stage
Operation
Control
Source freeze
Snapshot the six public benchmarks also evaluated by Graph-of-Agents.
Preserve source bytes, row order, split labels, counts, and hashes.
AI-assisted design
Use Astra to help develop and inspect the dynamic event-stream transformation.
Do not use model output to rewrite source tasks or establish gold answers.
Deterministic build
Map every source row to one episode; sample two within-dataset partners; generate 12 temporal queries.
Three reviewers inspect samples from all resulting dataset families.
Check source fidelity, temporal semantics, document–gold binding, and public/private separation.
Mechanical audit
Replay temporal selection with an independent validator and run regression tests.
Verify 31,119 episodes, 373,428 queries, schemas, manifests, and artifact checksums.
Appendix
Table 7: Dataset curation and quality-control stages. AI assistance supports protocol development, while deterministic generation and independent replay define the released artifacts and their gold bindings.
#
Public update before the query
a
w
Checkpoint
Selected
MMLU — mmlu-task-version-000000 . Documents: A = row 3370 (commander analogy); B = row 11280 (new-van liability); C = row 0 (field-extension degree).
1
Load B, A, and C; bind requests; activate C
2
2
Baseline
C
2
Supersede the active binding with B
7
7
Changed
B
3
Add an unrelated request; query the archive
12
12
Preservation
C
4
No new update; query the earlier world state
12
2
Historical
C
5
Schedule C for e=28
18
18
Preservation
B
Appendix
Table 8: Six actual episodes.
Dataset
Local split
Source items / episodes
Queries
Episode share
MMLU
test
14,042
168,504
45.1%
MMLU-Pro
test
12,032
144,384
38.7%
MedMCQA
test
4,183
50,196
13.4%
MATH
test
500
6,000
1.6%
GPQA
test
198
2,376
0.6%
HumanEval
dev
164
1,968
0.5%
Appendix
Table 9: The dynamic task routing collection. Every source item anchors one episode and every episode contains 12 checkpoint queries.
Dataset
Frozen export
Task field
Valid answer
Evaluator
MMLU
test/MMLU_test.json
question
A–D
Exact source option letter
MMLU-Pro
test/MMLU_Pro_test.json
question
A–J
Exact source option letter
MedMCQA
test/MedMCQA_test.json
question
A–D
Exact source option letter
GPQA
test/GPQA_test.json
question
A–D
Exact source option letter
MATH
test/MATH_test.json
question
String
Exact match to gold_answer
HumanEval
dev/human_eval_dev.json
prompt
Python
Hidden test with original entry_point
Appendix
Table 10: Dataset-specific input and output contracts. The public document is the indicated source field without rewriting. All final predictions also carry a document label in document_id .
Figure 9: Protocol-design comparison over the same source inventory. The top-left panel counts documents, checkpoint queries, event records, and public stream records per item or episode. Context lengths are medians of whitespace-separated visible tokens. The capability matrix states what is explicitly probed by the protocol, not measured model performance.
Figure 10: Additional diagnostic signal from Dynamic task routing. Left: one isolated response becomes 12 categorized checkpoints. Right: fraction of exact-answer episodes whose baseline and changed routes select different documents but share an identical gold answer string.
Figure 11: Semantic coverage of the 31,119 unique source tasks. (a) UMAP projection colored by source dataset. Axes and long-range projected distances have no direct semantic interpretation. (b) Dataset composition of each task’s 15 nearest neighbors, computed in the original normalized embedding space.
Figure 12: Mean pairwise cosine distance among the three source documents in each observed episode, compared with a within-dataset random-partner control. Distances are measured in the original normalized embedding space; boxes summarize episodes within each source dataset.
Figure 13: Descriptive semantic outcome landscapes for complete-run artifacts. (a) Homogeneous RIAG joint accuracy. (b) Heterogeneous minus homogeneous RIAG joint accuracy. (c) Homogeneous GoA mean minus max pooling. (d) Homogeneous GoA-mean joint accuracy. Colors are 50-neighbor averages in the original embedding space displayed on the common UMAP projection; they are not independent tests or causal effects.
Figure 14: Selection and joint accuracy by quartile of episode document separation, from closest (Q1) to farthest (Q4). Distance is the episode’s mean pairwise cosine distance in the original embedding space. Curves are descriptive aggregates over different mixtures of source datasets and do not identify a distance effect.
Q
Public update immediately before query
a
w
Request
Category
Gold pair
1
Load A, B, C; bind archive to C; set active task to C
6
6
Delivery
Baseline
(C,C)
2
Supersede the C assertion with active task B
8
8
Delivery
Changed
(B,D)
3
Add an unrelated request; query the fixed archive
10
10
Archive
Preservation
(C,C)
4
No update; query the delivery binding at world time 6
10
6
Delivery
Historical
(C,C)
5
Schedule active task A for effective time 21
14
14
Delivery
Preservation
(B,D)
6
No update; scheduled A is now effective
23
23
Delivery
Propagation
(A,A)
Appendix
Table 11: Complete context for the shared-episode example. Panel A gives all 12 queries in stream order; an update shown on a row arrives immediately before that query. Panel B gives the six recorded responses to highlighted query Q2 (ID c9e557d88cb3c4c275699e23 ). Correctness requires the complete (document, answer) pair, not a matching option symbol alone.