StateTree: Enhancing Long-Term Dialogue Reasoning via Reinforcement Learning
Organizations: Zhejiang University · Taobao & Tmall Group of Alibaba
Abstract
Large language models deployed as personalized assistants must reason over long, evolving interaction histories. However, in long-term dialogue reasoning, relevant evidence is scattered across sessions, preferences may be revised over time, and standard long-context training fails to address these challenges under data scarcity and prohibitive computational costs. We propose StateTree, a data-driven RL method that constructs a challenging auxiliary task from scarce dialogues with verifiable ground truth. StateTree augments multi-session dialogues with a tree-structured path-tracing task: key-value records are embedded across sessions to form a binary tree. Solving the task requires the model to traverse from root to leaf by retrieving records across sessions and comparing timestamps to resolve branches, then recover the hidden target question among distractor leaves. We apply curriculum RL training progressively increasing tree depth and introduce a compositional variant whose edges carry step-level reasoning fragments, training the model to compose partial cues into coherent queries. Trained on 10K-token contexts, StateTree generalizes to 128K tokens without full-length RL costs and exhibits capabilities including cross-session retrieval, temporal reasoning, knowledge update, and compositional multi-hop reasoning. StateTree outperforms both SFT and RL-based baselines while preserving short-context general reasoning. StateTree-7B achieves gains up to +23.60% on LongMemEval (128k), and StateTree-14B reaches 59.00% accuracy on LongMemEval, surpassing QwenLong-L1-32B (45.20%).
Figures & tables
| Challenge | Design in StateTree |
| Cross-Session Retrieval | Edge records distributed evenly across sessions, preventing positional shortcuts |
| Multi-Hop Reasoning | -level tree traversal; compositional StateTree requires aggregating step fragments into the question |
| Temporal Reasoning | At each fork, compare session timestamps to determine chronological ordering |
| Knowledge Update | Correct edge is placed in a newer session than the distractor, mirroring real recency preference |
| Model | Multi Hop | Temporal | Open Domain | Single Hop | Average | ||||||||||
| ACC | F1 | BLEU | ACC | F1 | BLEU | ACC | F1 | BLEU | ACC | F1 | BLEU | ACC | F1 | BLEU | |
| GPT-4o | 68.84 | 32.25 | 23.10 | 14.47 | 10.25 | 10.45 | 30.00 | 18.47 | 16.80 | 64.54 | 49.34 | 40.19 | 52.73 | 36.20 | 29.47 |
| QwenLong-L1-32B | 64.49 | 29.53 | 22.27 | 34.59 | 24.35 | 20.19 | 48.00 | 29.43 | 27.02 | 44.68 | 30.37 | 26.20 | 46.36 | 28.92 | 24.31 |
| R1-Distill-Qwen-32B | 70.29 | 26.80 | 17.45 | 33.33 | 15.98 | 12.35 | 52.00 | 25.19 | 20.98 | 52.01 | 23.04 | 18.00 | 51.43 | 22.40 | 16.93 |
| Qwen2.5-7B-Instruct | 63.77 | 21.34 | 15.23 | 23.90 | 12.67 | 10.58 | 50.00 | 16.88 | 14.08 | 44.92 | 21.96 | 17.72 | 44.29 | 19.60 | 15.56 |
| SEALONG-7B | 63.77 | 27.50 | 20.65 | 20.75 | 15.16 | 11.72 | 48.00 | 25.80 | 23.53 | 47.99 | 32.01 | 28.38 | 45.19 | 27.32 | 23.24 |
| LongMemEval (128k) | PersonaMem (128k) | ||||||||||||||
| Model | Temp. | Multi- Ses. | Know. Upd. | SS- User | SS- Asst. | SS- Pref. | Avg. | Latest Pref. | New Scen. | Align. Rec. | Shared Fact | Revisit Reas. | New Idea | Track Evol. | Avg. |
| QwenLong-L1-32B | 47.37 | 35.34 | 56.41 | 80.00 | 67.86 | 23.33 | 51.00 | 57.51 | 48.83 | 47.56 | 66.08 | 81.04 | 25.48 | 58.06 | 52.40 |
| R1-Distill-Qwen-32B | 21.80 | 13.53 | 50.00 | 34.29 | 57.14 | 10.00 | 29.00 | 33.03 | 21.60 | 33.52 | 36.84 | 63.57 | 22.39 | 66.57 | 37.62 |
| Qwen2.5-7B-Instruct | 20.30 | 8.27 | 52.56 | 25.71 | 37.50 | 3.33 | 23.80 | 37.41 | 36.15 | 45.56 | 47.95 | 62.08 | 15.64 | 62.76 | 40.48 |
| SEALONG-7B | 28.57 | 28.57 | 64.10 | 75.71 | 76.79 | 16.67 | 45.40 | 37.99 | 39.44 | 46.99 | 47.37 | 64.68 | 16.80 | 63.64 | 41.66 |
| LoongRL-7B | 24.06 | 5.26 | 55.13 | 37.14 | 42.86 | 3.33 | 26.60 | 29.56 | 23.00 | 39.83 | 35.09 | 41.26 | 21.81 | 37.54 | 31.39 |
| Model | LoCoMo | PersonaMem | LongMemEval | |
| 10K | 32K | 128K | 128K | |
| Qwen2.5-7B-Instruct | 44.29 | 53.14 | 40.48 | 23.80 |
| StateTree-7B | 49.61 +5.32 | 61.29 +8.15 | 47.30 +6.82 | 47.40 +23.60 |
| Qwen2.5-14B-Instruct | 49.09 | 58.91 | 45.40 | 45.20 |
| StateTree-14B | 60.91 +11.82 | 64.52 +5.61 | 52.59 +7.19 | 59.00 +13.80 |
| Qwen3-8B | 43.64 | 58.23 | 45.36 | 49.20 |
| Model | MMLU | MATH | IFEval | Avg. |
| Qwen2.5-7B-Instruct | 73.4 | 76.0 | 71.2 | 73.5 |
| StateTree-7B | 73.5 +0.10 | 76.0 +0.00 | 70.5 -0.70 | 73.3 -0.20 |
| Qwen2.5-14B-Instruct | 79.4 | 83.4 | 81.0 | 81.3 |
| StateTree-14B | 79.8 +0.40 | 83.0 -0.40 | 79.3 -1.70 | 80.7 -0.60 |
| Qwen3-8B | 76.9 | 78.2 | 85.0 | 80.0 |
| StateTree-8B | 77.0 +0.10 | 78.1 -0.10 | 85.2 +0.20 | 80.1 +0.10 |
| Variant | LoCoMo | LongMemEval | PersonaMem |
| StateTree-7B (full) | 49.61 | 47.40 | 47.30 |
| w/o warm-up | 45.19 -4.42 | 43.00 -4.40 | 44.00 -3.30 |
| w/o Stage 1 ( ) | 44.68 -4.93 | 42.60 -4.80 | 43.78 -3.52 |
| w/o Stage 2 ( ) | 46.23 -3.38 | 44.20 -3.20 | 45.10 -2.20 |
| w/o Stage 3 (Compositional) | 46.75 -2.86 | 44.80 -2.60 | 45.51 -1.79 |
| Variant | LoCoMo | LongMemEval | PersonaMem |
| Qwen2.5-7B-Instruct | 44.29 | 23.80 | 40.48 |
| w/ CoT prompt | 45.19 +0.90 | 24.60 +0.80 | 41.11 +0.63 |
| StateTree-7B | 49.61 | 47.40 | 47.30 |
| w/ Independent Random Keys | 45.45 -4.16 | 25.20 -22.20 | 41.03 -6.27 |
| w/ Single Session | 47.79 -1.82 | 42.80 -4.60 | 43.96 -3.34 |
| w/ Entity Keys | 48.44 -1.17 | 45.60 -1.80 | 45.98 -1.32 |
| Reward | LoCoMo | LongMemEval | PersonaMem |
| Combined Reward | 49.61 | 47.40 | 47.30 |
| Exact Match only | 47.79 -1.82 | 46.80 -0.60 | 46.68 -0.62 |
| LLM-as-a-Judge only | 48.70 -0.91 | 46.60 -0.80 | 46.75 -0.55 |
| Token-level F1 | 47.14 -2.47 | 45.00 -2.40 | 45.29 -2.01 |
| Two-way Substr. EM | 44.68 -4.93 | 42.40 -5.00 | 43.31 -3.99 |
| ROUGE-L | 42.60 -7.01 | 39.40 -8.00 | 41.29 -6.01 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | Warm-up | Stage 1 | Stage 2 | Stage 3 |
| Tree type | None (direct QA) | Basic StateTree | Basic StateTree | Compositional StateTree |
| Tree depth | – | 2 | 3 | 3 |
| Leaves | – | 4 | 8 | 8 |
| Edge records | 0 | 6 | 14 | 14 |
| Samples | 616 | 616 | 616 | 616 |
| Avg. length (tokens) | 10K | 10K | 10K | 10K |
| Level 1 | Level 2 (Event) | Level 3 (Question) |
| [A] Jon | [B] expanding his studio’s social media presence | When did [A] start [B]? |
| What did [A] do while [B]? | ||
| [B] going to a fair for exposure | When did [A] start [B]? | |
| What events did [A] participate in while [B]? | ||
| [A] Gina | [B] launching an ad campaign | When did [A] start [B]? |
| Why did [A] decide to start [B]? |
| garden | mirror | temple | anchor | lantern | ribbon | candle |
| shield | bridge | castle | meadow | feather | saddle | trumpet |
| Model | Generalizing to New Scenarios | Provide Preference Aligned Recommendations | Recall User Shared Facts | Recalling Facts Mentioned by the User | Recalling the Reasons Behind Previous Updates | Suggest New Ideas | Track Full Preference Evolution | Average |
| QwenLong-L1-32B | 71.93 | 65.45 | 67.44 | 52.94 | 78.79 | 22.58 | 74.82 | 63.84 |
| R1-Distill-Qwen-32B | 56.14 | 61.82 | 61.24 | 64.71 | 83.84 | 22.58 | 79.14 | 62.82 |
| Qwen2.5-7B-Instruct | 52.63 | 60.00 | 41.86 | 52.94 | 79.80 | 17.20 | 66.19 | 53.14 |
| SEALONG-7B | 56.14 | 54.55 | 41.09 | 47.06 | 82.83 | 17.20 | 64.75 | 52.80 |
| LoongRL-7B | 59.65 | 69.09 | 48.84 | 64.71 | 78.79 | 20.43 | 72.66 | 58.40 |
| RL-MemAgent-7B | 52.63 | 49.09 | 47.29 | 76.47 | 79.80 | 34.41 | 57.55 | 54.67 |
| Variant | Multi Hop | Temporal | Open Domain | Single Hop | Average |
| Qwen2.5-7B-Instruct | 63.77 | 23.90 | 50.00 | 44.92 | 44.29 |
| StateTree-7B (full) | 64.49 | 32.70 | 56.00 | 50.35 | 49.61 |
| w/o warm-up | 62.32 | 27.67 | 50.00 | 45.63 | 45.19 |
| w/o Stage 1 ( ) | 63.04 | 24.53 | 50.00 | 45.63 | 44.68 |
| w/o Stage 2 ( ) | 60.14 | 28.93 | 52.00 | 47.52 | 46.23 |
| w/o Stage 3 (Comp.) | 60.87 | 29.56 | 52.00 | 47.99 | 46.75 |
| Variant | Temporal | Multi-Session | Knowledge Update | Single-Session User | Single-Session Assistant | Single-Session Preference | Average |
| Qwen2.5-7B-Instruct | 20.30 | 8.27 | 52.56 | 25.71 | 37.50 | 3.33 | 23.80 |
| StateTree-7B (full) | 29.32 | 31.58 | 67.95 | 75.71 | 76.79 | 23.33 | 47.40 |
| w/o warm-up | 27.82 | 28.57 | 64.10 | 68.57 | 67.86 | 13.33 | 43.00 |
| w/o Stage 1 ( ) | 24.81 | 25.56 | 62.82 | 71.43 | 73.21 | 20.00 | 42.60 |
| w/o Stage 2 ( ) | 26.32 | 28.57 | 64.10 | 72.86 | 73.21 | 20.00 | 44.20 |
| w/o Stage 3 (Comp.) | 27.07 | 28.57 | 65.38 | 72.86 | 75.00 | 20.00 | 44.80 |
| Level 1 (Root) | Level 2 (Next) | Level 3 (Question) |
| 387c1508-38e1-43f0- a36d-ce8cba77a4c9 | 37cbcf3a-944f-45cd- b647-e5e12f51d593 | What is Caroline’s identity? |
| What was the poetry reading that Caroline attended about? | ||
| 77376180-3363-4fa3- 95d7-f21859335c9e | What does Melanie do to destress? | |
| When did Melanie’s friend adopt a child? |
| Level 1 | Level 2 (Event) | Level 3 (Question / Answer) |
| [A] Caroline | [B] discussing her identity (correct) | What is [A]’s identity while [B]? Transgender woman |
| What career paths is [A] considering while [B]? Psychology, counseling certification | ||
| [B] researching adoption agencies | What did [A] learn while [B]? Adoption agencies | |
| When did [A] start [B]? researching adoption agencies | ||
| [A] Melanie | [B] running a charity race | When did [A] participate in [B]? The sunday before 25 May 2023 |
| What did [B] that [A] was part of raise awareness for? mental health |