MemStrata: 95% and 90.91% Source-Aware Accuracy on LongMemEval-500 and LoCoMo-1540 with a Local Qwen 3.8 27B Q4_K_M Reader
Organizations: Called It Inc.
Abstract
An adequate conversational answer may differ from a short or incomplete benchmark reference. To measure adequacy against the recorded history we prefer source-aware grading, in which the judge checks the reference against the full source before assessing system-blinded answers; original reference-only grading is reported alongside. With a local Qwen 3.8 27B Q4_K_M reader and a 24,000-token evidence ceiling, MemStrata CL1 scores 475/500 (95.0%) on LongMemEval-S and 1,400/1,540 (90.91%) on LoCoMo categories 1-4 under source-aware GPT-5.5 adjudication, against 463/500 (92.6%) and 1,205/1,540 (78.25%) under reference-only grading of the same answers. It preserves a retrieval backbone and adds nonduplicated, dated, speaker-attributed source spans. A same-reader full-history control with about 4.7 times the evidence scores 464/500 reference-only and 470/500 (94.0%) source-aware; neither difference is decisive. Keyword-only selection at the same budget scores 425, and a matched-reader Letta arm 438. On LongMemEval-M, where the packet holds about 1.6% of each history, MemStrata CL1 scores 427/500, with losses concentrated in multi-session and temporal questions. On 300 BEAM-1M questions it outscores dense retrieval, 0.738 to 0.706 (Wilcoxon p = 0.011). A same-seed replay of unchanged requests changed 1.5-2.3% of labels. On identical packets GLM 5.3 flash is non-inferior within 3 points (462 versus 463); Muse Spark 1.3 did not show non-inferiority on 269 questions. None of four pre-registered interventions met all of its registered advancement or feasibility criteria. Signed read-side artifacts support inspection but do not regenerate the private retrieval pipeline. The superiority of source-aware grading to human adjudication is not established, and development exposure, automated-judge dependence and the absence of held-out data preclude an independent-replication or leaderboard claim.
Figures & tables
| Configuration | Correct/ | Accuracy | Comparison status |
|---|---|---|---|
| MemStrata CL1 , 24k | 463/500 | 92.6% | Primary |
| Letta evidence, 100k ceiling | 438/500 | 87.6% | Matched reader |
| MemStrata 7k | 420/500 | 84.0% | Historical |
| Letta evidence, 8k ceiling | 387/500 | 77.4% | Historical |
| Question category | MemStrata 7k | Letta 8k | Letta 100k | MemStrata CL1 | |
|---|---|---|---|---|---|
| Single-session user | 70 | 67 (95.7) | 60 (85.7) | 63 (90.0) | 66 (94.3) |
| Single-session assistant | 56 | 56 (100.0) | 56 (100.0) | 56 (100.0) | 56 (100.0) |
| Single-session preference | 30 | 20 (66.7) | 26 (86.7) | 25 (83.3) | 26 (86.7) |
| Multi-session | 133 | 92 (69.2) | 87 (65.4) | 107 (80.5) | 117 (88.0) |
| Knowledge update | 78 | 73 (93.6) | 63 (80.8) | 72 (92.3) | 74 (94.9) |
| Temporal reasoning | 133 | 112 (84.2) | 95 (71.4) | 115 (86.5) | 124 (93.2) |
| Evidence configuration | Correct/ | Accuracy |
|---|---|---|
| MemStrata 7k, historical R230 | 351/500 | 70.2% |
| Letta archival 8k | 289/500 | 57.8% |
| Letta archival 100k ceiling | 237/500 | 47.4% |
| Configuration | Correct/ | Accuracy | Status |
|---|---|---|---|
| MemStrata CL1 , 24k | 1,205/1,540 | 78.25% | Primary; 4 failure zeros |
| MemStrata baseline, 8,192 cap | 1,122/1,540 | 72.86% | Historical; 2 failure zeros |
| Letta evidence + Qwen | 1,020/1,540 | 66.23% | Retained comparator |
| Frozen Zep evidence + Qwen | 913/1,540 | 59.29% | Retained comparator |
| Category | Baseline | MemStrata CL1 | Change (pp) | |
|---|---|---|---|---|
| Multi-hop (1) | 282 | 113 (40.07) | 159 (56.38) | +16.31 |
| Temporal (2) | 321 | 244 (76.01) | 261 (81.31) | +5.30 |
| Open-domain (3) | 96 | 41 (42.71) | 45 (46.88) | +4.17 |
| Single-hop (4) | 841 | 724 (86.09) | 740 (87.99) | +1.90 |
| Overall | 1,540 | 1,122 (72.86) | 1,205 (78.25) | +5.39 |
| Evidence configuration | Mean tokens | Median tokens | Maximum |
|---|---|---|---|
| LongMemEval MemStrata CL1 | 23,335 | 23,970 | 24,000 |
| LongMemEval Letta 100k ceiling | 23,878 | 24,036 | 29,011 |
| LoCoMo MemStrata CL1 | 20,857 | – | |
| LoCoMo historical MemStrata | 5,485 | – | 8,192 |
| Configuration | Reference-only | Source-aware | Up | Down |
|---|---|---|---|---|
| MemStrata CL1 24k | 463/500 (92.6%) | 475/500 (95.0%) | 24 | 12 |
| Letta, 100k ceiling | 438/500 (87.6%) | 444/500 (88.8%) | 20 | 14 |
| Category | MemStrata CL1 reference | MemStrata CL1 source | Letta source | |
|---|---|---|---|---|
| Single-session user | 70 | 66 | 69 (98.57%) | 65 (92.86%) |
| Single-session assistant | 56 | 56 | 56 (100.00%) | 56 (100.00%) |
| Single-session preference | 30 | 26 | 29 (96.67%) | 25 (83.33%) |
| Multi-session | 133 | 117 | 124 (93.23%) | 113 (84.96%) |
| Knowledge update | 78 | 74 | 74 (94.87%) | 73 (93.59%) |
| Temporal reasoning | 133 | 124 | 123 (92.48%) | 112 (84.21%) |
| Configuration | Reference-only | Source-aware | Up | Down |
|---|---|---|---|---|
| MemStrata CL1 24k | 1,205 (78.25%) | 1,400 (90.91%) | 211 | 16 |
| MemStrata 8k | 1,122 (72.86%) | 1,291 (83.83%) | 185 | 16 |
| Letta 8k | 1,020 (66.23%) | 1,185 (76.95%) | 181 | 16 |
| Frozen Zep 8k | 913 (59.29%) | 963 (62.53%) | 158 | 108 |
| Category | MemStrata CL1 | MemStrata 8k | Letta 8k | Zep 8k | |
|---|---|---|---|---|---|
| Multi-hop | 282 | 159 / 231 56.38 / 81.91 | 113 / 160 40.07 / 56.74 | 117 / 171 41.49 / 60.64 | 112 / 147 39.72 / 52.13 |
| Temporal | 321 | 261 / 302 81.31 / 94.08 | 244 / 286 76.01 / 89.10 | 229 / 263 71.34 / 81.93 | 178 / 136 55.45 / 42.37 |
| Open-domain | 96 | 45 / 74 46.88 / 77.08 | 41 / 70 42.71 / 72.92 | 37 / 71 38.54 / 73.96 | 41 / 71 42.71 / 73.96 |
| Single-hop | 841 | 740 / 793 87.99 / 94.29 | 724 / 775 86.09 / 92.15 | 637 / 680 75.74 / 80.86 | 582 / 609 69.20 / 72.41 |
| Arm | Source-aware | Downgrades | Upper bound | Upper bound (%) |
| LongMemEval MemStrata CL1 , 24k | 475 | 12 | 487/500 | 97.40 |
| LongMemEval Letta, 100k ceiling | 444 | 14 | 458/500 | 91.60 |
| LoCoMo MemStrata CL1 , 24k | 1,400 | 16 | 1,416/1,540 | 91.95 |
| LoCoMo MemStrata 8k | 1,291 | 16 | 1,307/1,540 | 84.87 |
| LoCoMo Letta 8k | 1,185 | 16 | 1,201/1,540 | 77.99 |
| LoCoMo frozen Zep 8k | 963 | 108 | 1,071/1,540 | 69.55 |
| System | Original OD /96 | Knowledge OD /96 | Overall, old updated |
|---|---|---|---|
| MemStrata CL1 24k | 45 (46.88%) | 53 (55.21%) | 78.25 78.77% |
| MemStrata 8k | 41 (42.71%) | 48 (50.00%) | 72.86 73.31% |
| Letta 8k | 37 (38.54%) | 46 (47.92%) | 66.23 66.82% |
| Frozen Zep 8k | 41 (42.71%) | 47 (48.96%) | 59.29 59.68% |
| System | Prior OD /96 | New OD /96 | Gains / losses | Combined overall |
|---|---|---|---|---|
| MemStrata CL1 24k | 74 (77.08%) | 82 (85.42%) | 15 / 7 | 1,408 (91.43%) |
| MemStrata 8k | 70 (72.92%) | 75 (78.13%) | 10 / 5 | 1,296 (84.16%) |
| Letta 8k | 71 (73.96%) | 73 (76.04%) | 7 / 5 | 1,187 (77.08%) |
| Frozen Zep 8k | 71 (73.96%) | 74 (77.08%) | 10 / 7 | 966 (62.73%) |
| Fresh-reader replay quantity | Result |
| Questions / newly generated reader answers | 500 / 500 |
| Byte-identical answers / reused original grades | 467 / 467 |
| Changed answers / fresh GPT-5.5 grades | 33 / 33 |
| Primary reference correctness | 463/500 (92.6%) |
| Fresh-reader correctness | 464/500 (92.8%) |
| Incorrect-to-correct / correct-to-incorrect flips | 2 / 1 |
| Evidence on LongMemEval-S | Mean tokens | Reference-only | Arm-only / CL1-only | Exact | Source-aware |
|---|---|---|---|---|---|
| MemStrata CL1 , 24k ceiling | 23,335 | 463 (92.6%) | – | – | 475 (95.0%) |
| Full history, untrimmed | 108,556 | 464 (92.8%) | 21 / 20 | 1.0 | 470 (94.0%) |
| Full history, as registered (35 trimmed) | – | 466 (93.2%) | 21 / 18 | 0.75 | 472 (94.4%) |
| NOBACKBONE, 24k ceiling | 22,514 | 457 (91.4%) | 13 / 19 | 0.38 | – |
| DENSE, 24k ceiling | 23,977 | 451 (90.2%) | 12 / 24 | 0.065 | – |
| BM25, 24k ceiling | 20,522 | 425 (85.0%) | 11 / 49 | – |
| Category (same questions) | LongMemEval-S | LongMemEval-M | Gained / lost | Exact | |
|---|---|---|---|---|---|
| Single-session user | 70 | 66 | 67 | 2 / 1 | 1.0 |
| Single-session assistant | 56 | 56 | 56 | 0 / 0 | 1.0 |
| Single-session preference | 30 | 26 | 21 | 3 / 8 | 0.23 |
| Multi-session | 133 | 117 | 102 | 3 / 18 | 0.0015 |
| Knowledge update | 78 | 74 | 71 | 2 / 5 | 0.45 |
| Temporal reasoning | 133 | 124 | 110 | 4 / 18 | 0.0043 |
| BEAM-1M question type ( each) | MemStrata CL1 | DENSE | Difference (points) |
|---|---|---|---|
| Temporal reasoning | 0.692 | 0.608 | +8.3 |
| Instruction following | 0.917 | 0.864 | +5.3 |
| Multi-session reasoning | 0.762 | 0.715 | +4.7 |
| Contradiction resolution | 0.788 | 0.742 | +4.6 |
| Abstention | 0.575 | 0.533 | +4.2 |
| Information extraction | 0.864 | 0.842 | +2.3 |
| Reader on identical LongMemEval-S MemStrata CL1 packets | Reader | Qwen | Status |
|---|---|---|---|
| GLM 5.3 flash, maximum reasoning (500) | 462 (92.4%) | 463 (92.6%) | Confirmatory: non-inferior within 3 points |
| Muse Spark 1.3, extra-high reasoning (500) | 449 (89.8%) | 463 (92.6%) | Descriptive; registered 269-question test: non-inferiority not shown |
| Study | Intervention | Primary result | Verdict |
|---|---|---|---|
| Output limit | Reader output limit raised from 16,384 to 32,768 tokens | Accuracy within the guards on all four sets (LongMemEval-S 183 vs 185 of 200; LoCoMo 312 vs 311 of 400; LongMemEval-M 179 vs 179 of 200; BEAM 0.7355 vs 0.7378). Output failures against the unchanged replay: LongMemEval-M 0 vs 2, BEAM 1 vs 4 | Did not meet the output-failure criterion: 4 LoCoMo output failures against 2 in the anchor (tolerance +1) |
| Reader instructions | Revised reader prompt (conflict handling, brevity) | Pooled primary 493/600 vs 490/600 (15 gains, 12 losses, ). LongMemEval-S 187 vs 185; LoCoMo 318 vs 311; LongMemEval-M 175 vs 179; BEAM 0.7326 vs 0.7378 | Did not meet the output-failure criterion (LongMemEval-M 3 vs 1); pooled gain not significant |
| Fallback call | One extra non-thinking call after an empty or length-limited answer (26 failed answers) | 21/26 valid completions; 4/14 text answers correct; median added time 65.9 s against a 60 s limit. The variant without the failed reasoning: 24/26 valid, 4/14 correct, median 60.3 s | Did not meet text quality (4/14 vs 7/14), BEAM quality (0.275 vs 0.50) or latency (65.9 s vs 60 s); yield passed |
| Breadth answers | Round-robin admission across time strata for detected summarization and event-ordering questions on BEAM; two replicates per arm | Mean judged-score change over 60 registered targets (25 gains, 23 losses and 10 ties among 58 detected; sign-flip ); output failures 6 vs 4 | Did not meet the gain criterion or the output-failure guard (limit 5) |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Field | Recorded value |
|---|---|
| Endpoint / model tag | /api/chat ; qwen3.8:27b |
| Quantization | Q4_K_M, recorded local artifact |
| Evidence ceiling | 24,000 tokens, complete rendered evidence |
| num_ctx / num_predict | 128000 / 16384 |
| temperature / top_p / top_k | 1.0 / 0.95 / 20 |
| min_p / repeat_penalty / presence_penalty | 0.0 / 1.0 / 0.0 |
| Configuration | Correct/ | Accuracy | Gains/losses vs MemStrata CL1 | Exact | Abstention |
|---|---|---|---|---|---|
| MemStrata CL1 (primary) | 463/500 | 92.6% | – | – | 28/30 |
| Adaptive temperature | 463/500 | 92.6% | 6/6 | 1.0 | 28/30 |
| R | 460/500 | 92.0% | 10/13 | 0.678 | 27/30 |
| CS1 | 459/500 | 91.8% | 10/14 | 0.541 | 28/30 |
| LoCoMo MemStrata CL1 (primary) | 1,205/1,540 | 78.25% | – | – | – |
| LoCoMo CS1 | 1,203/1,540 | 78.12% | 45/47 | 0.917 | – |
| LongMemEval-S category | MemStrata CL1 | Adaptive | R | CS1 | |
|---|---|---|---|---|---|
| Single-session user | 70 | 66 | 66 | 67 | 66 |
| Single-session assistant | 56 | 56 | 56 | 56 | 56 |
| Single-session preference | 30 | 26 | 25 | 26 | 25 |
| Multi-session | 133 | 117 | 119 | 113 | 113 |
| Knowledge update | 78 | 74 | 74 | 73 | 75 |
| Temporal reasoning | 133 | 124 | 123 | 125 | 124 |
| Diagnostic condition | Correct/32 | Source-aware | Gain / loss vs. prior |
|---|---|---|---|
| A: retained original answers | 23 | 71.88% | – |
| B: general-knowledge permission | 26 | 81.25% | 4 / 1 |
| C: knowledge + original images | 25 | 78.13% | 2 / 3 |
| D: images + conditional web | 20 | 62.50% | 1 / 6 |
| Artifact | Identity |
|---|---|
| LongMemEval-S cleaned dataset | d6f21ea9d60a0d56f34a05b609c79c88a451d2ae03597821ea3d5a9678c3a442 |
| LoCoMo locomo10.json | 79fa87e90f04081343b8c8debecb80a9a6842b76a7aa537dc9fdf651ea698ff4 |
| LongMemEval-M cleaned dataset | 9d79e5524794a2e6900a3aa9cb7d9152c5a3e8319c9a87c25494ba1eacee495f |
| BEAM-1M prepared questions | 393377d727c080552527901bb2791f6386fb765f8625badb4d4d9c06b948fb04 |
| Reader-substitution study seal (GLM) | 0b8a58665272ea916b1e7f9679fec624dd63f60b38dcf5f6042123611fb204a2 |
| Reader-substitution study seal (Muse) | d01a39ccbf0a64a4489a0465e8c7da1ccc7bb4210a94e0c3cf2a69a255395467 |