While Large Language Models (LLMs) advertise million-token context windows, reasoning quality often collapses as inputs grow -- a phenomenon termed context rot. This failure stems from a structural entanglement in monolithic architectures, where the massive search burden of contextual grounding exhausts the representational capacity needed for complex reasoning. To resolve this, we propose Grounding-Reasoning Disaggregation via DIStributed long COntext scaling (DISCO). Inspired by distributed computing frameworks like Apache Spark, DISCO partitions long context across a fleet of Worker LLMs dedicated exclusively to parallel, localized grounding. A central Driver LLM, trained via Reinforcement Learning (GRPO) to optimize planning, orchestrates execution by dynamically mapping queries into atomic extraction tasks and reducing the gathered evidence to synthesize a final answer. By isolating reasoning from raw context noise, DISCO effectively eliminates context rot. On RULER-QA (1M tokens), it maintains 78.4% accuracy where standard baselines collapse. Furthermore, it outperforms full-context models by up to 9.8 points on LongBench v2 and matches frontier models like Gemini-3-Pro-Preview while reducing inference costs by over 80%, establishing a highly efficient paradigm for robust long-context inference.
Figures & tables
Figure 1: The execution pipeline of DISCO. A Driver LLM orchestrates the inference by dynamically generating a DAG of actions: Map (Narrow) tasks are offloaded to parallel Worker LLMs for evidence extraction, while Reduce (Wide) tasks are handled by the Driver for global synthesis and reasoning. This cycle repeats iteratively until the query is resolved.
LongBench v2
RULER-QA
∞ Bench
Method
Medium
Long
Avg
256K
512K
1M
Avg
En.MC
En.QA
Avg
Qwen3-8B (Thinking Mode)
Full Long Context
28.8
32.4
30.6
—
—
—
—
65.94
49.86
57.90
RAG
30.8
32.1
31.5
52.4
33.5
10.9
32.27
66.35
53.15
59.75
CoA
29.4
24.0
26.7
44.8
46.0
44.6
45.13
41.96
22.38
32.17
LLM × MapReduce
31.9
29.2
30.6
75.3
73.2
72.4
73.63
48.03
42.16
45.10
Table 1: Evaluation results of models on long-context benchmarks. LongBench v2 and RULER-QA numbers are accuracy. For ∞ Bench , we report MC (Multiple Choice), QA (Question Answering), and their average accuracy. DISCO (RL) represents our proposed method after GRPO training. See Table 4 for more model families results.
Configuration
Acc (%)
Avg Stages
Latency (s)
Full Reward
39.8
3.39
381
w/o Revd
36.3
4.45
432
w/o Rfmt
34.4
3.84
499
Table 2: Ablation of Reward Design on LongBench v2 (Qwen3-8B). Latency includes retries from format errors.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
LongBench v2
Worker Type
Medium
Long
Avg
Qwen3-Embedding-4B
37.2
41.9
39.6
Qwen3-4B-Instruct
43.7
50.9
47.30
Appendix
Table 3: Comparison of DISCO using a Generative Worker versus a Dense Retriever baseline. The Driver is Qwen3-14B after our RL training.
Driver
Worker
Score
LLaMA-3.1-70B-Instruct
— (Full Context, CoT)
25.90
Ministral-3-8B-Instruct (FP8)
28.70
Qwen3-4B-Instruct
27.80
Ministral-8B-Reasoning
— (Full Context, CoT)
32.41
Qwen3-4B-Instruct-2507
37.96
Gemini-3-Pro-Preview
Ministral-3-8B-Instruct (FP8)
70.37
Appendix
Table 4: Generalization of DISCO across diverse Driver and Worker model combinations on LongBench v2 (Long) . DISCO consistently improves over the Full-Context CoT baseline regardless of the driver model family.
Understanding and reasoning over long contexts has become a key requirement for deploying large language models (LLMs) in realistic applications. Although recent LLMs support increasingly long context windows, they often fail to use relevant evidence that is already present in the input, revealing a gap between context access and effective context utilization. In this work, we propose Recursive Evidence Replay as LLM Harness for Long-Context Reasoning (RECONTEXT), a training-free inference method for improving long-context reasoning. RECONTEXT uses model-internal relevance signals to construct a query-conditioned evidence pool and replays it before final generation while preserving the full original context. This recursive selection process separates evidence organization from answer generation without training, external memory, or context pruning. We also provide a theoretical analysis based on associative memory, which characterizes the context as a memory store, the question as a retrieval cue, attention as cue-trace association, and replay as trace reactivation. Experiments on eight long-context datasets with 128K context length show that RECONTEXT consistently improves evidence utilization across Qwen3-4B, Qwen3-8B, and Llama3-8B, achieving the best average rank on all three backbones. Code is available at https://github.com/Yanjun-Zhao/ReContext.
Long-context reasoning requires accurately identifying relevant information in extensive, noisy input contexts. In this work, we propose PERK (Parameter Efficient Reasoning over Knowledge), a scalable approach for learning to encode long contexts using gradient updates at test time. Specifically, PERK employs two nested optimization loops in a meta-training phase. The inner loop rapidly encodes contexts into a low-rank adapter (LoRA) that serves as a parameter-efficient memory module for the base model. Concurrently, the outer loop learns to use the updated adapter to accurately recall and reason over relevant information from the encoded long context. Our evaluations on several long-context reasoning tasks show that PERK significantly outperforms the standard long-context finetuning, achieving average absolute performance gains of up to 20% for Qwen-2.5 (0.5B & 7B) on synthetic and real-world long-context reasoning. PERK also maintains its advantages across model scales and families. Compared to specialized long-context LLMs, PERK matches or surpasses their performance. Finally, our analyses show PERK is more robust to reasoning complexity, length extrapolation, and the positions of relevant information in contexts. https://perk-long-context.web.app
Long-context reasoning remains a critical bottleneck for large language models, as recent recurrent-memory approaches face two inherent challenges: sequential chunk-wise updates can overwrite early critical evidence with later irrelevant content, and serial inter-chunk dependencies limit parallelism and cause latency to increase with context length. To address these issues, we propose PI-Mem (Parallel-Iterative Memory), a mechanism that processes all chunks in parallel and iteratively refines a shared memory over a bounded number of turns. In each turn, PI-Mem reads all chunks in parallel conditioned on the current memory, selects new or complementary evidence from each chunk, and merges the selected evidence into a compact shared memory for the next turn. To discourage redundant turns, we optimize the workflow through reinforcement learning with an auxiliary turn-efficiency reward, enabling the model to adaptively exit once sufficient evidence has been accumulated. We evaluate PI-Mem with Qwen3.5-35B-A3B and Qwen2.5-7B on the HotpotQA benchmark across context lengths up to 3.6 million tokens and find that it outperforms the recurrent-memory baseline by +6.25 and +7.81 absolute points while achieving 6.1× and 2.1× inference speedups, respectively. These results demonstrate that PI-Mem breaks the accuracy--efficiency trade-off in long-context reasoning and provides a scalable approach to complex multi-hop question answering over extremely long documents.
Dawei Liu, Haixu Song, Shuang Cheng +9
Shanghai Jiao Tong University · Shanghai Artificial Intelligence Laboratory · Tsinghua University +3