Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context
Organizations: Department of Electrical and Computer Engineering University of California San Diego
Abstract
Conversations with chat assistants increasingly span many topics in a single long-running thread, challenging memory systems. Existing long-context and memory benchmarks often expose session or topic boundaries, or probe direct personal-memory questions. These settings understate a harder assistant-memory regime: a flat mixed-topic thread where the system must infer which earlier episode makes a later task decision valid. We introduce SCALE-QA, a constraint-grounded task QA benchmark for flat unsegmented threads targeting episode integrity failure. The dataset contains 3,000 audited questions across 10 domains, uses deterministic four-way multiple-choice grading, and includes a deterministic runtime builder; experiments use all 3,000 questions through 128k and a stratified 400-question diagnostic at 1M. SCALE-QA questions are ordinary task-oriented requests whose correct answer depends on causally related evidence introduced earlier in the conversation. We also propose Temporal-Semantic Interleaved Memory Reconstruction (TSIM), which segments the turn stream into coherent episodes and indexes them through a hierarchical multi-view memory stack with deterministic episode-level summary and cluster-routing views. Experiments show that SCALE-QA challenges strong RAG baselines and long-context LLMs alike; across three open-source and proprietary LLM backends, TSIM achieves the highest accuracy in every backend setting, gaining 5.6-17.6 accuracy points over the strongest corresponding baseline.
Figures & tables
| Failure mode | Failure pattern | Bottleneck | Recovery unit |
|---|---|---|---|
| Retrieval miss | evidence not retrieved | visibility | passage |
| Lost-in-the-middle | evidence underused in prompt | position | token span |
| State tracking failure | updates not maintained | update order | state value |
| Episode integrity failure | evidence visible but fragmented | episode integrity | operative episode |
| Backend | Retriever | Acc | CL Hit | Rat. Sim. | CtxTok | Lat. |
|---|---|---|---|---|---|---|
| Gemma2:9b | TSIM | 69.6 | 70.7 | 0.472 | 1060.8 | 4.20 |
| Standard RAG | 24.4 | 7.7 | 0.296 | 929.5 | 3.35 | |
| Hybrid-RRF Chunk RAG | 31.1 | 11.2 | 0.254 | 1193.6 | 4.00 | |
| RAPTOR | 40.6 | 35.3 | 0.403 | 790.5 | 20.67 | |
| MemGPT | 60.1 | 62.6 | 0.413 | 2301.9 | 3.88 | |
| HippoRAG | 25.2 | 12.5 | 0.380 | 748.9 | 6.51 |
| Stage | Variant | Acc | CL | Rat. | Lat. | Tok. |
|---|---|---|---|---|---|---|
| L0 | Std. RAG top- | 26.2 | 5.6 | 0.246 | 1.61 | 968.4 |
| L1 | Fixed-token direct | 43.4 | 35.0 | 0.416 | 3.52 | 1269.2 |
| L2 | Semantic-drift direct | 55.5 | 52.2 | 0.430 | 3.62 | 1179.8 |
| L3 | Full TSIM stack | 74.2 | 74.3 | 0.485 | 3.56 | 1044.4 |
Appendix figures & tables28 assets
Supplementary material from the paper’s appendix.
Appendix
| Axis | LongMemEval | SCALE-QA |
|---|---|---|
| History format | Timestamped chat histories with session structure | Flat mixed-topic threads with no boundary metadata |
| Question type | Flexible personal-memory QA | Constraint-grounded task QA |
| Domain ontology | Personal-life ontology (health, hobbies, work-life, etc.) | Ten task-oriented operational domains |
| Target failure mode | Long-term memory over sessions and updates | Episode integrity failure in unsegmented threads |
| Evaluation protocol | Judged open-ended responses | Deterministic four-way MCQ with exact evidence traces |
| Domain | QA | Typical evidence forms | Common retrieval stress cues |
|---|---|---|---|
| CS-Software | 300 | code/version/deployment rules | stale defaults; tool exceptions |
| Network/Hardware | 300 | port, routing, hardware policies | constraint traps; local compliance |
| Finance | 300 | credit, reimbursement, risk rules | overwritten eligibility; denials |
| Legal | 300 | contracts, filings, exceptions | exception resolution; operative clauses |
| Biomed | 300 | protocols, safety notes, exclusions | dormant safety constraints |
| Engineering | 300 | materials, tests, device specs | incompatibilities; hidden failures |
| Release property | Value |
|---|---|
| Total QA records | 3,000 |
| Public split labels | 2 1,500 |
| Topics | 10 300 |
| Split-topic cells | 20 150 |
| Correct labels | A/B/C/D = 750/750/750/750 |
| Audited evidence snippets | 4,346 |
| Split | QA | Dom. | Evid. | Zero | Oracle |
|---|---|---|---|---|---|
| Full dataset | 3,000 | 3,000/3,000 | 7.63/18.10 | 100.00/100.00 | |
| Codex | 1,500 | 1,500/1,500 | 7.40/19.27 | 100.00/100.00 | |
| Claude | 1,500 | 1,500/1,500 | 7.87/16.93 | 100.00/100.00 |
| Metric | Mean | Median | Std. | Direction |
|---|---|---|---|---|
| Naturalness | 3.80 | 3.67 | 0.25 | higher better |
| Answerability | 4.91 | 5.00 | 0.17 | higher better |
| Ambiguity risk | 1.45 | 1.33 | 0.32 | lower better |
| Constraint plausibility | 3.99 | 4.00 | 0.29 | higher better |
| Ann. | Gold Acc. | Max ans. | Natural | Answerable | Ambig. | Plausible |
|---|---|---|---|---|---|---|
| Ann. 1 | 99.0 | 25.7 | 4.17 | 4.90 | 1.19 | 3.98 |
| Ann. 2 | 91.0 | 30.3 | 4.04 | 4.87 | 1.92 | 4.08 |
| Ann. 3 | 96.3 | 25.3 | 3.20 | 4.95 | 1.25 | 3.90 |
| Pair | Ans. agr. | Ans. | Primary agr. | State | Long | Trap | Cons. gold |
|---|---|---|---|---|---|---|---|
| Ann. 1–2 | 91.3 | 0.884 | 56.7 | 43.3 | 83.0 | 84.7 | 272/274 |
| Ann. 1–3 | 96.0 | 0.947 | 16.3 | 45.3 | 83.0 | 100.0 | 287/288 |
| Ann. 2–3 | 89.0 | 0.853 | 32.7 | 57.3 | 72.0 | 84.7 | 265/267 |
| View | Label | Count | Rate |
|---|---|---|---|
| Multi-label cue | State overwrite | 196 | 65.3 |
| Multi-label cue | Long-range bridge | 291 | 97.0 |
| Multi-label cue | Constraint trap | 300 | 100.0 |
| Multi-label cue | Any multi-label | 298 | 99.3 |
| Multi-label cue | All three cues | 189 | 63.0 |
| Primary cue | Constraint trap | 158 | 52.7 |
| Topic | Natural | Answerable | Ambig. | Plausible |
|---|---|---|---|---|
| Biomed | 3.66 | 4.86 | 1.49 | 4.00 |
| Biz-Ops | 3.79 | 4.92 | 1.31 | 4.13 |
| CS-Software | 3.74 | 4.94 | 1.34 | 4.09 |
| Daily-Life | 4.09 | 4.91 | 1.38 | 4.04 |
| Engineering | 3.72 | 4.90 | 1.42 | 4.13 |
| Finance | 3.79 | 4.88 | 1.41 | 4.06 |
| Protocol item | Definition |
|---|---|
| Token accounting | Benchmark-side estimate whitespace word count, used for deterministic packing and constructed-length reporting, not for asserting exact provider-native token parity. |
| Length scaling | No built-in length cap; practical limits are distractor-pool size and computational budget. This paper reports settings through 1M to match evaluated native-context backends. |
| Full-corpus regime | Used when selected truth tokens fit inside the target context length; the full selected truth background is retained and noise fills remaining space. |
| Length-batched regime | Used when selected truth tokens exceed the target length; records are deterministically partitioned into budget-controlled batches. |
| Packing rule | Deterministic capacity-constrained best fit: records are ordered by descending truth-token count and then by question name. |
| Truth cap | In length-batched mode, the default truth-cap ratio is , leaving room for noise and prompt/interface overhead. |
| System | Retrieval unit | Context assembly | Purpose in comparison |
|---|---|---|---|
| Standard RAG top-5 | Individual retrieved chunks/turns | Top five retrieved units are passed to the answer backend. | Tests whether short semantic retrieval alone can recover the decisive evidence. |
| Hybrid-RRF Chunk RAG | Dense chunks + BM25 sparse hits + neighboring turns | Dense and sparse hits are fused with reciprocal-rank fusion, expanded with local neighbors, and packed under a matched context cap. | Tests whether lightweight hybrid chunk retrieval closes the episode-reconstruction gap. |
| Tuned Hybrid-Rerank Chunk RAG | Dense/sparse candidates + HyDE + cross-encoder reranking | Adds query expansion, RRF fusion, BGE reranking, parent-window expansion, and MMR packing under the same GPT-4o-mini 128k task. | Tests whether a strongly tuned but non-episodic chunk pipeline can close the gap, and exposes its reranking cost. |
| RAPTOR strict no-block | Hierarchical summaries built without gold blocks | Retrieved hierarchy outputs are mapped into the same no-block evaluation regime. | Tests whether hierarchical abstraction alone solves interleaved conversational memory. |
| MemGPT paper-default | Explicit memory-management substrate | Retrieved memory material is inserted under the same answer and writeback protocol. | Tests whether higher recall from a memory-style substrate converts into final accuracy. |
| HippoRAG paper-default | Graph-style retrieval substrate | Retrieved graph/memory evidence is evaluated with the same question set and scoring protocol. | Tests transfer of graph-centric memory retrieval to flat-thread writeback evaluation. |
| Streaming semantic-drift segmentation in Official TSIM |
|---|
| 1. Flatten the mixed conversation into a turn stream and encode each turn as a normalized dense vector. |
| 2. Maintain the current segment start and token count . |
| 3. For incoming turn , average only the recent turns still inside the current segment to form local center . |
| 4. Score the newest turn by cosine similarity to , plus a small bonus for a model user transition. |
| 5. If and the score falls below similarity threshold , close the segment before turn . |
| 6. If , force a boundary even if semantic drift is weak. |
| Stage | Data | Purpose |
|---|---|---|
| Search | 300 QA | Fast screening of candidate retrieval configurations. |
| Confirm | 600 QA | Check stability across 64k, 128k, 256k, and 512k targets. |
| Answer validation | 200 QA | Confirm that retrieval gains transfer to answer accuracy. |
| Full retrieval check | 3,000 QA | Verify evidence coverage and context stability at 128k. |
| Candidate | Representative variation | Search Hit | Confirm Hit | Answer Hit | Answer Acc | Answer Ctx |
|---|---|---|---|---|---|---|
| Official calibrated | 86.33 | 79.33 | 87.50 | 78.50 | 1011.6 | |
| High-threshold compact | 85.67 | 77.62 | 86.00 | 79.00 | 757.7 | |
| Lean memory stack | 84.00 | 76.58 | 87.50 | 77.50 | 818.2 | |
| Lean L2-exp | Same as lean, one L2-to-episode expansion | 83.67 | 76.85 | 87.50 | 77.50 | 818.1 |
| Component | Frozen reported constants |
|---|---|
| Segmenter | min_tokens =120, max_tokens =320, recent_window =4, , speaker bonus , minimum two-message buffer. |
| Retrieval breadth | , , , . |
| Ranking weights | , , , , . |
| Expansion | Cluster expansion , L2-to-episode expansion , cluster threshold , soft margin , temporal expansion hops . |
| Prompt assembly | Reconstructed episodes capped at eight messages with two-message overlap; L2 cluster summaries route and boost candidates but are not inserted directly into the prompt. |
| Gemma2:9b | ||||||
|---|---|---|---|---|---|---|
| Domain | TSIM | Std. RAG | Hybrid | RAPTOR | MemGPT | HippoRAG |
| CS | 79.0 | 29.3 | 33.7 | 48.0 | 66.7 | 25.3 |
| Net | 77.3 | 26.7 | 33.7 | 49.3 | 67.0 | 27.0 |
| Fin | 54.3 | 22.7 | 29.3 | 32.7 | 46.7 | 23.3 |
| Legal | 60.0 | 22.3 | 28.0 | 36.3 | 56.3 | 26.7 |
| Bio | 74.7 | 25.3 | 36.7 | 39.7 | 66.7 | 21.3 |
| Stress view | Std. RAG | Hybrid-RRF | RAPTOR | MemGPT | HippoRAG | TSIM | |
|---|---|---|---|---|---|---|---|
| All audited | 300 | 26.7/5.7 | 29.0/9.3 | 41.7/29.7 | 52.7/58.3 | 29.3/8.7 | 73.7/78.7 |
| State overwrite | 196 | 24.0/3.6 | 27.0/8.2 | 38.3/26.5 | 49.5/56.6 | 28.6/7.1 | 70.9/76.0 |
| Long-range bridge | 291 | 26.8/5.8 | 28.9/8.9 | 42.3/29.9 | 52.9/58.1 | 29.6/8.6 | 73.9/78.4 |
| Constraint trap | 300 | 26.7/5.7 | 29.0/9.3 | 41.7/29.7 | 52.7/58.3 | 29.3/8.7 | 73.7/78.7 |
| Retrieved unit | All recall@5 | Co-contain |
|---|---|---|
| TSIM episode | 0.810 | 0.890 |
| Fixed 128-token window | 0.719 | 0.675 |
| Fixed 256-token window | 0.647 | 0.826 |
| Fixed 320-token window | 0.577 | 0.831 |
| Standard RAG chunk | 0.456 | 0.004 |
| Method | Embedder | All recall@5 |
|---|---|---|
| Standard RAG | BGE-large | 0.456 |
| Standard RAG | MiniLM | 0.408 |
| TSIM | BGE-large | 0.810 |
| TSIM | MiniLM | 0.649 |
| Method | Boundary information | Judged Acc. (%) | All-evidence recall (%) | Avg. answer- context tokens |
|---|---|---|---|---|
| TSIM | None; episodes reconstructed | 71.0 | 84.04 | 4,723.5 |
| BM25 fixed-chunk top-5 | None | 61.2 | 65.74 | 4,616.3 |
| BGE turn top-5 | None | 56.6 | 70.64 | 340.8 |
| BM25 session top-5 | Official session boundaries | 74.4 | 76.60 | 13,086.5 |
| System | Views | Vectors | Vector MB | Text amp. | Ingest ms/turn | Ret. p50 ms | Ret. p95 ms | Summary API |
|---|---|---|---|---|---|---|---|---|
| Standard RAG | 1 | 5,554 | 21.7 | 1.056 | 4.178 | 21.478 | 36.791 | 0 |
| TSIM | 3 | 6,281 | 24.5 | 2.087 | 24.986 | 141.093 | 159.978 | 0 |
| Failure type | Example | Diagnostic observation |
|---|---|---|
| Episode-incomplete retrieval | codex:Network-Hardware-075 | Standard RAG retrieves related text but omits the neighboring license clause that makes the local decision operative. |
| Fixed-window low SNR | claude-code:Network-Hardware-032 | A fixed window contains the evidence within 1,755 tokens but at lower evidence density than the 1,127-token reconstructed episode. |
| Answer-side override misuse | claude-code:Biomed-003 | TSIM retrieves the decisive episode, yet the answer model follows a plausible public default instead of the local override. |
| Routing/packing miss | codex:Biomed-124 | Two gold units are required; final selected context covers only one, exposing a residual routing and packing failure. |
| Budget | Hit | Acc | HitRank | Lat. | CtxTok |
|---|---|---|---|---|---|
| 16k | 100.00 | 62.5 | 335.4 | 4.32 | 29,938 |
| 32k | 100.00 | 57.8 | 669.2 | 4.57 | 60,145 |
| 64k | 100.00 | 41.7 | 836.3 | 10.03 | 86,315 |
| 128k | 100.00 | 29.8 | 711.7 | 14.61 | 88,425 |
| System | Prompt | Rows | Acc. | Prompt tok. | Lat. |
|---|---|---|---|---|---|
| Full Context | Vanilla | 100 | 49.0 | 98,339 | 5.52 |
| Full Context | Dev-selected | 100 | 49.0 | 98,298 | 2.80 |
| TSIM | – | 100 | 77.0 | 1,389 | 2.40 |
| Backend | Method | Acc | Hit-C | Hit-W | Miss-C | Miss-W |
|---|---|---|---|---|---|---|
| Gemma2:9b | Standard RAG | 24.4 | 5.2 | 2.1 | 10.6 | 47.7 |
| Hybrid-RRF Chunk RAG | 31.1 | 8.4 | 2.1 | 9.5 | 27.2 | |
| RAPTOR | 40.6 | 25.2 | 10.0 | 14.6 | 47.0 | |
| MemGPT | 60.1 | 48.6 | 14.0 | 11.4 | 25.9 | |
| HippoRAG | 25.2 | 8.7 | 3.7 | 15.6 | 68.4 | |
| TSIM | 69.6 | 58.2 | 12.4 | 11.3 | 17.7 |
| Backend | Baseline | Base | TSIM | 95% CI | |
|---|---|---|---|---|---|
| Gemma2:9b | Standard RAG | 24.4 | 69.6 | +45.21 | [43.48, 47.03] |
| Hybrid-RRF Chunk RAG | 31.1 | 69.6 | +38.52 | [36.72, 40.24] | |
| RAPTOR | 40.6 | 69.6 | +28.98 | [27.13, 30.76] | |
| MemGPT | 60.1 | 69.6 | +9.53 | [7.90, 11.22] | |
| HippoRAG | 25.2 | 69.6 | +44.42 | [42.41, 46.36] | |
| Gemini 2.5 Flash | Standard RAG | 29.8 | 80.2 | +50.33 | [48.98, 51.63] |
| Method | Domain Acc Std. | Batch Acc Range | Ctx Mean | Ctx Median | Ctx P95 | Ctx Max |
|---|---|---|---|---|---|---|
| Standard RAG | 3.73 | 20.4–30.0 | 929.5 | 937 | 1076 | 1259 |
| Hybrid-RRF Chunk RAG | 3.65 | 25.0–43.2 | 1193.6 | 1194 | 1200 | 1200 |
| RAPTOR | 7.62 | 37.4–45.5 | 790.5 | 780 | 996 | 1219 |
| MemGPT | 7.70 | 52.6–72.7 | 2301.9 | 2317 | 2608 | 3359 |
| HippoRAG | 3.97 | 20.8–29.7 | 748.9 | 757 | 891 | 1043 |
| TSIM | 8.08 | 64.5–76.9 | 1060.8 | 1037 | 1408 | 1950 |