While Large Language Models (LLMs) advertise million-token context windows, reasoning quality often collapses as inputs grow -- a phenomenon termed context rot. This failure stems from a structural entanglement in monolithic architectures, where the massive search burden of contextual grounding exhausts the representational capacity needed for complex reasoning. To resolve this, we propose Grounding-Reasoning Disaggregation via DIStributed long COntext scaling (DISCO). Inspired by distributed computing frameworks like Apache Spark, DISCO partitions long context across a fleet of Worker LLMs dedicated exclusively to parallel, localized grounding. A central Driver LLM, trained via Reinforcement Learning (GRPO) to optimize planning, orchestrates execution by dynamically mapping queries into atomic extraction tasks and reducing the gathered evidence to synthesize a final answer. By isolating reasoning from raw context noise, DISCO effectively eliminates context rot. On RULER-QA (1M tokens), it maintains 78.4% accuracy where standard baselines collapse. Furthermore, it outperforms full-context models by up to 9.8 points on LongBench v2 and matches frontier models like Gemini-3-Pro-Preview while reducing inference costs by over 80%, establishing a highly efficient paradigm for robust long-context inference.
Figures & tables
Figure 1: The execution pipeline of DISCO. A Driver LLM orchestrates the inference by dynamically generating a DAG of actions: Map (Narrow) tasks are offloaded to parallel Worker LLMs for evidence extraction, while Reduce (Wide) tasks are handled by the Driver for global synthesis and reasoning. This cycle repeats iteratively until the query is resolved.
LongBench v2
RULER-QA
∞ Bench
Method
Medium
Long
Avg
256K
512K
1M
Avg
En.MC
En.QA
Avg
Qwen3-8B (Thinking Mode)
Full Long Context
28.8
32.4
30.6
—
—
—
—
65.94
49.86
57.90
RAG
30.8
32.1
31.5
52.4
33.5
10.9
32.27
66.35
53.15
59.75
CoA
29.4
24.0
26.7
44.8
46.0
44.6
45.13
41.96
22.38
32.17
LLM × MapReduce
31.9
29.2
30.6
75.3
73.2
72.4
73.63
48.03
42.16
45.10
Table 1: Evaluation results of models on long-context benchmarks. LongBench v2 and RULER-QA numbers are accuracy. For ∞ Bench , we report MC (Multiple Choice), QA (Question Answering), and their average accuracy. DISCO (RL) represents our proposed method after GRPO training. See Table 4 for more model families results.
Configuration
Acc (%)
Avg Stages
Latency (s)
Full Reward
39.8
3.39
381
w/o Revd
36.3
4.45
432
w/o Rfmt
34.4
3.84
499
Table 2: Ablation of Reward Design on LongBench v2 (Qwen3-8B). Latency includes retries from format errors.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
LongBench v2
Worker Type
Medium
Long
Avg
Qwen3-Embedding-4B
37.2
41.9
39.6
Qwen3-4B-Instruct
43.7
50.9
47.30
Appendix
Table 3: Comparison of DISCO using a Generative Worker versus a Dense Retriever baseline. The Driver is Qwen3-14B after our RL training.
Driver
Worker
Score
LLaMA-3.1-70B-Instruct
— (Full Context, CoT)
25.90
Ministral-3-8B-Instruct (FP8)
28.70
Qwen3-4B-Instruct
27.80
Ministral-8B-Reasoning
— (Full Context, CoT)
32.41
Qwen3-4B-Instruct-2507
37.96
Gemini-3-Pro-Preview
Ministral-3-8B-Instruct (FP8)
70.37
Appendix
Table 4: Generalization of DISCO across diverse Driver and Worker model combinations on LongBench v2 (Long) . DISCO consistently improves over the Full-Context CoT baseline regardless of the driver model family.