Large language models deployed as personalized assistants must reason over long, evolving interaction histories. However, in long-term dialogue reasoning, relevant evidence is scattered across sessions, preferences may be revised over time, and standard long-context training fails to address these challenges under data scarcity and prohibitive computational costs. We propose StateTree, a data-driven RL method that constructs a challenging auxiliary task from scarce dialogues with verifiable ground truth. StateTree augments multi-session dialogues with a tree-structured path-tracing task: key-value records are embedded across sessions to form a binary tree. Solving the task requires the model to traverse from root to leaf by retrieving records across sessions and comparing timestamps to resolve branches, then recover the hidden target question among distractor leaves. We apply curriculum RL training progressively increasing tree depth and introduce a compositional variant whose edges carry step-level reasoning fragments, training the model to compose partial cues into coherent queries. Trained on 10K-token contexts, StateTree generalizes to 128K tokens without full-length RL costs and exhibits capabilities including cross-session retrieval, temporal reasoning, knowledge update, and compositional multi-hop reasoning. StateTree outperforms both SFT and RL-based baselines while preserving short-context general reasoning. StateTree-7B achieves gains up to +23.60% on LongMemEval (128k), and StateTree-14B reaches 59.00% accuracy on LongMemEval, surpassing QwenLong-L1-32B (45.20%).
Figures & tables
Figure 1: Model trajectories in multi-session dialogue reasoning. (i) QwenLong-L1-32B gets lost across sessions and misses information. (ii) The StateTree-trained model shows cross-session retrieval, temporal reasoning, knowledge update, and multi-hop reasoning, improving reasoning reliability.
Figure 2: Examples of the challenges in multi-turn dialogue. For each example, we show the associated evidence statements on the left and the question with the answer on the right.
Figure 3: The StateTree data construction pipeline. Left: Basic StateTree embeds key-value records across sessions into a tree; the model traverses root-to-leaf to recover the target question. Right: Compositional StateTree augments edges with step-level fragments aggregates into the final question.
Challenge
Design in StateTree
Cross-Session Retrieval
Edge records distributed evenly across S sessions, preventing positional shortcuts
Multi-Hop Reasoning
D -level tree traversal; compositional StateTree requires aggregating step fragments into the question
Temporal Reasoning
At each fork, compare session timestamps to determine chronological ordering
Knowledge Update
Correct edge is placed in a newer session than the distractor, mirroring real recency preference
Table 1: Mapping between multi-turn dialogue challenges and StateTree design choices.
Model
Multi Hop
Temporal
Open Domain
Single Hop
Average
ACC
F1
BLEU
ACC
F1
BLEU
ACC
F1
BLEU
ACC
F1
BLEU
ACC
F1
BLEU
GPT-4o
68.84
32.25
23.10
14.47
10.25
10.45
30.00
18.47
16.80
64.54
49.34
40.19
52.73
36.20
29.47
QwenLong-L1-32B
64.49
29.53
22.27
34.59
24.35
20.19
48.00
29.43
27.02
44.68
30.37
26.20
46.36
28.92
24.31
R1-Distill-Qwen-32B
70.29
26.80
17.45
33.33
15.98
12.35
52.00
25.19
20.98
52.01
23.04
18.00
51.43
22.40
16.93
Qwen2.5-7B-Instruct
63.77
21.34
15.23
23.90
12.67
10.58
50.00
16.88
14.08
44.92
21.96
17.72
44.29
19.60
15.56
SEALONG-7B
63.77
27.50
20.65
20.75
15.16
11.72
48.00
25.80
23.53
47.99
32.01
28.38
45.19
27.32
23.24
Table 2: Main results on the LoCoMo benchmark across different task dimensions.
LongMemEval (128k)
PersonaMem (128k)
Model
Temp.
Multi- Ses.
Know. Upd.
SS- User
SS- Asst.
SS- Pref.
Avg.
Latest Pref.
New Scen.
Align. Rec.
Shared Fact
Revisit Reas.
New Idea
Track Evol.
Avg.
QwenLong-L1-32B
47.37
35.34
56.41
80.00
67.86
23.33
51.00
57.51
48.83
47.56
66.08
81.04
25.48
58.06
52.40
R1-Distill-Qwen-32B
21.80
13.53
50.00
34.29
57.14
10.00
29.00
33.03
21.60
33.52
36.84
63.57
22.39
66.57
37.62
Qwen2.5-7B-Instruct
20.30
8.27
52.56
25.71
37.50
3.33
23.80
37.41
36.15
45.56
47.95
62.08
15.64
62.76
40.48
SEALONG-7B
28.57
28.57
64.10
75.71
76.79
16.67
45.40
37.99
39.44
46.99
47.37
64.68
16.80
63.64
41.66
LoongRL-7B
24.06
5.26
55.13
37.14
42.86
3.33
26.60
29.56
23.00
39.83
35.09
41.26
21.81
37.54
31.39
Table 3: OOD generalization accuracy (%) on LongMemEval and PersonaMem-128k benchmarks.
Model
LoCoMo
PersonaMem
LongMemEval
10K
32K
128K
128K
Qwen2.5-7B-Instruct
44.29
53.14
40.48
23.80
StateTree-7B
49.61 +5.32
61.29 +8.15
47.30 +6.82
47.40 +23.60
Qwen2.5-14B-Instruct
49.09
58.91
45.40
45.20
StateTree-14B
60.91 +11.82
64.52 +5.61
52.59 +7.19
59.00 +13.80
Qwen3-8B
43.64
58.23
45.36
49.20
Table 4: Average accuracy (%) across context lengths.
Model
MMLU
MATH
IFEval
Avg.
Qwen2.5-7B-Instruct
73.4
76.0
71.2
73.5
StateTree-7B
73.5 +0.10
76.0 +0.00
70.5 -0.70
73.3 -0.20
Qwen2.5-14B-Instruct
79.4
83.4
81.0
81.3
StateTree-14B
79.8 +0.40
83.0 -0.40
79.3 -1.70
80.7 -0.60
Qwen3-8B
76.9
78.2
85.0
80.0
StateTree-8B
77.0 +0.10
78.1 -0.10
85.2 +0.20
80.1 +0.10
Table 5: Short-context reasoning performance (%).
Figure 4: Reasoning F1 score and response lengths throughout RL training on the validation set.
Variant
LoCoMo
LongMemEval
PersonaMem
StateTree-7B (full)
49.61
47.40
47.30
w/o warm-up
45.19 -4.42
43.00 -4.40
44.00 -3.30
w/o Stage 1 ( D=2 )
44.68 -4.93
42.60 -4.80
43.78 -3.52
w/o Stage 2 ( D=3 )
46.23 -3.38
44.20 -3.20
45.10 -2.20
w/o Stage 3 (Compositional)
46.75 -2.86
44.80 -2.60
45.51 -1.79
Table 6: Curriculum stage ablation.
Variant
LoCoMo
LongMemEval
PersonaMem
Qwen2.5-7B-Instruct
44.29
23.80
40.48
w/ CoT prompt
45.19 +0.90
24.60 +0.80
41.11 +0.63
StateTree-7B
49.61
47.40
47.30
w/ Independent Random Keys
45.45 -4.16
25.20 -22.20
41.03 -6.27
w/ Single Session
47.79 -1.82
42.80 -4.60
43.96 -3.34
w/ Entity Keys
48.44 -1.17
45.60 -1.80
45.98 -1.32
Table 7: Temporal discrimination ablation.
Reward
LoCoMo
LongMemEval
PersonaMem
Combined Reward
49.61
47.40
47.30
Exact Match only
47.79 -1.82
46.80 -0.60
46.68 -0.62
LLM-as-a-Judge only
48.70 -0.91
46.60 -0.80
46.75 -0.55
Token-level F1
47.14 -2.47
45.00 -2.40
45.29 -2.01
Two-way Substr. EM
44.68 -4.93
42.40 -5.00
43.31 -3.99
ROUGE-L
42.60 -7.01
39.40 -8.00
41.29 -6.01
Table 8: Reward function ablation.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Warm-up
Stage 1
Stage 2
Stage 3
Tree type
None (direct QA)
Basic StateTree
Basic StateTree
Compositional StateTree
Tree depth D
–
2
3
3
Leaves
–
4
8
8
Edge records
0
6
14
14
Samples
616
616
616
616
Avg. length (tokens)
∼ 10K
∼ 10K
∼ 10K
∼ 10K
Appendix
Table 9: Dataset statistics and data construction parameters for each curriculum stage. “UUID format” refers to the identifier format used for tree node keys. “Insert mode” specifies where records are placed within the dialogue text.
Level 1
Level 2 (Event)
Level 3 (Question)
[A] Jon
[B] expanding his studio’s social media presence
★ When did [A] start [B]?
What did [A] do while [B]?
[B] going to a fair for exposure
When did [A] start [B]?
What events did [A] participate in while [B]?
[A] Gina
[B] launching an ad campaign
When did [A] start [B]?
Why did [A] decide to start [B]?
Appendix
Table 10: Semantic decomposition example for the target question “ When did Jon start expanding his studio’s social media presence? ” (answer: “April, 2023”). ★ marks the target leaf. The correct path aggregates: step_1 = “[A] Jon”, step_2 = “[B] expanding his studio’s social media presence”, step_3 = “When did [A] start [B]?”, yielding the full question after substitution.
garden
mirror
temple
anchor
lantern
ribbon
candle
shield
bridge
castle
meadow
feather
saddle
trumpet
Appendix
Table 11: Entity key pool used in the w/ Entity Keys ablation.
Model
Generalizing to New Scenarios
Provide Preference Aligned Recommendations
Recall User Shared Facts
Recalling Facts Mentioned by the User
Recalling the Reasons Behind Previous Updates
Suggest New Ideas
Track Full Preference Evolution
Average
QwenLong-L1-32B
71.93
65.45
67.44
52.94
78.79
22.58
74.82
63.84
R1-Distill-Qwen-32B
56.14
61.82
61.24
64.71
83.84
22.58
79.14
62.82
Qwen2.5-7B-Instruct
52.63
60.00
41.86
52.94
79.80
17.20
66.19
53.14
SEALONG-7B
56.14
54.55
41.09
47.06
82.83
17.20
64.75
52.80
LoongRL-7B
59.65
69.09
48.84
64.71
78.79
20.43
72.66
58.40
RL-MemAgent-7B
52.63
49.09
47.29
76.47
79.80
34.41
57.55
54.67
Appendix
Table 12: Performance comparison across different models on OOD PersonaMem-32k.
Variant
Multi Hop
Temporal
Open Domain
Single Hop
Average
Qwen2.5-7B-Instruct
63.77
23.90
50.00
44.92
44.29
StateTree-7B (full)
64.49
32.70
56.00
50.35
49.61
w/o warm-up
62.32
27.67
50.00
45.63
45.19
w/o Stage 1 ( D=2 )
63.04
24.53
50.00
45.63
44.68
w/o Stage 2 ( D=3 )
60.14
28.93
52.00
47.52
46.23
w/o Stage 3 (Comp.)
60.87
29.56
52.00
47.99
46.75
Appendix
Table 13: Category-level curriculum ablation on LoCoMo (7B).
Variant
Temporal
Multi-Session
Knowledge Update
Single-Session User
Single-Session Assistant
Single-Session Preference
Average
Qwen2.5-7B-Instruct
20.30
8.27
52.56
25.71
37.50
3.33
23.80
StateTree-7B (full)
29.32
31.58
67.95
75.71
76.79
23.33
47.40
w/o warm-up
27.82
28.57
64.10
68.57
67.86
13.33
43.00
w/o Stage 1 ( D=2 )
24.81
25.56
62.82
71.43
73.21
20.00
42.60
w/o Stage 2 ( D=3 )
26.32
28.57
64.10
72.86
73.21
20.00
44.20
w/o Stage 3 (Comp.)
27.07
28.57
65.38
72.86
75.00
20.00
44.80
Appendix
Table 14: Category-level curriculum ablation on LongMemEval (7B).
Level 1 (Root)
Level 2 (Next)
Level 3 (Question)
387c1508-38e1-43f0- a36d-ce8cba77a4c9
37cbcf3a-944f-45cd- b647-e5e12f51d593
★ What is Caroline’s identity?
What was the poetry reading that Caroline attended about?
77376180-3363-4fa3- 95d7-f21859335c9e
What does Melanie do to destress?
When did Melanie’s friend adopt a child?
Appendix
Table 15: StateTree structure for the D=2 Basic StateTree sample (conv-26). ★ marks the target leaf.
Level 1
Level 2 (Event)
Level 3 (Question / Answer)
[A] Caroline
[B] discussing her identity (correct)
★ What is [A]’s identity while [B]? → Transgender woman
What career paths is [A] considering while [B]? → Psychology, counseling certification
[B] researching adoption agencies
What did [A] learn while [B]? → Adoption agencies
When did [A] start [B]? → researching adoption agencies
[A] Melanie
[B] running a charity race
When did [A] participate in [B]? → The sunday before 25 May 2023
What did [B] that [A] was part of raise awareness for? → mental health
Appendix
Table 16: Full 2×2×2 tree structure for the Compositional StateTree sample (conv-26). ★ marks the target leaf.
Understanding and reasoning over long contexts has become a key requirement for deploying large language models (LLMs) in realistic applications. Although recent LLMs support increasingly long context windows, they often fail to use relevant evidence that is already present in the input, revealing a gap between context access and effective context utilization. In this work, we propose Recursive Evidence Replay as LLM Harness for Long-Context Reasoning (RECONTEXT), a training-free inference method for improving long-context reasoning. RECONTEXT uses model-internal relevance signals to construct a query-conditioned evidence pool and replays it before final generation while preserving the full original context. This recursive selection process separates evidence organization from answer generation without training, external memory, or context pruning. We also provide a theoretical analysis based on associative memory, which characterizes the context as a memory store, the question as a retrieval cue, attention as cue-trace association, and replay as trace reactivation. Experiments on eight long-context datasets with 128K context length show that RECONTEXT consistently improves evidence utilization across Qwen3-4B, Qwen3-8B, and Llama3-8B, achieving the best average rank on all three backbones. Code is available at https://github.com/Yanjun-Zhao/ReContext.
Long-context reasoning is an essential capability for large language models, particularly when they are deployed as autonomous agents that must reason over lengthy trajectories. Reinforcement learning (RL) has recently emerged as a dominant paradigm for improving this ability, yet existing work largely focuses on reward engineering while diverse training data remains scarce. We revisit this problem from a data-centric perspective and show that a simple yet effective data recipe alone, paired with a minimal outcome-based GRPO setup, suffices to substantially improve long-context reasoning. Our recipe targets three complementary task families -- retrieval, multi-evidence synthesis, and reasoning -- for which we construct and curate eight datasets totaling ~14K examples. Experiments on three models (Qwen3-4B/8B/30B-A3B) yield average gains of +7.2/+3.2/+6.4 points across seven long-context benchmarks, surpassing prior RL training sets. We further demonstrate that these gains transfer to agentic tasks, where continuing RL training on an agent-tuned model with our data recipe improves GAIA by +4.8 and BrowseComp by +7.0 points. We will release our datasets to facilitate future research.
Long-context reasoning requires models to locate, revise, and synthesize evidence distributed across lengthy inputs. Existing long-context RL methods usually reward final answers or static evidence extraction, offering little feedback on how intermediate actions change the model's evidence state. We propose Maven, a reinforcement learning framework with an editable evidence memory. Maven defines an answer-conditioned evidence-state value and rewards action-level state transitions: add actions are credited by marginal gain and hindsight contribution, link actions by evidence synergy, and drop actions by improved answer support after removing misleading evidence. These rewards are assigned to the corresponding action spans in GRPO. Across Llama and Qwen models on LongBench v2, LongReason, and RULER, Maven outperforms outcome-only RL and evidence-identification baselines, producing more sufficient evidence sets and lower distractor retention. Our results show that long-context RL benefits from optimizing stateful evidence navigation rather than one-shot evidence extraction.