When Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision Model
Organizations: Independent Researcher
Abstract
Does conversational memory need LLM-extracted facts, or is selecting the right raw turns enough? Published results disagree. Extraction-based systems report gains from distilled facts. Recent studies find raw history with good ranking does as well, but disagree about whether ranking matters. We ran a pre-registered study on held-out LoCoMo conversations and LongMemEval. At a tight budget on LoCoMo, raw turns selected by a single call to Jev, a typed decision model, are non-inferior to an LLM-extraction memory (one-sided 95% bound -3.0 points against a -5-point margin). Blind human grading narrows the margin but does not change the result. Raw turns cost 3,061 times less to write, and the result holds with a second answer model. Within this study, reranking's gain shrinks as the budget grows. It adds 17.4 points on LoCoMo and 9.1 on LongMemEval when three of 30 candidates are kept. At generous budgets it adds 1.5 and 1.1, and extraction systems are more accurate. This suggests why published results disagree. At matched context, Jev selects as accurately as an LLM reranker (non-inferiority bound -2.0) at a third of the latency, and more accurately than a multi-call graph traversal. Reranking lowers correct abstention. Plans, code and graded answers are released.
Figures & tables
| Grading | Turns + Jev | engram v2 | Difference | One-sided 95% bound | Two-sided 95% CI | Non-inferior (margin 5) |
|---|---|---|---|---|---|---|
| Judge (registered) | 77.0% | 77.5% | 0.5 | 3.0 | [ 3.5, +2.5] | yes |
| Human, strict | 76.3% | 78.0% | 1.7 | 3.9 | [ 4.3, +0.9] | yes |
| Human, lenient | 77.4% | 79.9% | 2.6 | 4.7 | [ 5.1, 0.03] | yes |
| Test | Turns + Jev vs | Turns + Jev k | Turns + Jev | Other | Only Turns + Jev / only other | p | Holm p |
|---|---|---|---|---|---|---|---|
| S1 | Turns + cosine (LoCoMo) | 3 | 77.2% | 59.9% | 152 / 17 | 2.7e 28 | 1.9e 27 |
| S2 | Jev-Mem (LoCoMo) | 4 | 77.0% | 70.6% | 98 / 48 | 4.3e 5 | 1.3e 4 |
| S3 | mem0 (LoCoMo) | 3 | 77.2% | 68.5% | 134 / 66 | 1.7e 6 | 7.6e 6 |
| S4 | Turns + LLM (LoCoMo, non-inferiority) | 3 | 77.2% | 77.6% | 28 / 31 | 1.5e 6 | 7.6e 6 |
| S5 | mem0 (LongMemEval, 30 knowledge-update) | 2 | 70.0% | 70.0% | 4 / 4 | 1.00 | 1.00 |
| S6 | Turns + cosine (LongMemEval sample, 70) | 3 | 68.6% | 65.7% | 7 / 5 | 0.77 | 1.00 |
| System, setting | Accuracy | Tokens | Write $/1k | Read $/query | Multi-hop | Temporal | Open-domain | Single-hop |
| Tight budget ( 3 or 6) | ||||||||
| Turns + cosine, 3 | 59.9% | 130 | $0.0006 | 0 † | 54.3 | 55.2 | 42.0 | 65.7 |
| Turns + cosine, 6 | 68.6% | 254 | $0.0006 | 0 † | 61.4 | 60.6 | 54.0 | 75.9 |
| Turns + Jev, 3 | 77.2% | 139 | $0.0006 | $0.00020 | 75.7 | 69.1 | 58.0 | 83.2 |
| Turns + Jev, 6 | 77.0% | 265 | $0.0006 | $0.00020 | 74.3 | 72.1 | 56.0 | 82.3 |
| Turns + LLM, 3 | 77.6% | 143 | $0.0006 | $0.00023 | 73.6 | 72.7 | 58.0 | 83.2 |
| System | Accuracy | Tokens | KU (72) | MS (121) | SS-A (56) | SS-P (30) | SS-U (64) | TR (127) | Abs |
| Tight budget ( 3) | |||||||||
| Turns + cosine, 3 | 57.7% | 522 | 58.3 | 31.4 | 85.7 | 53.3 | 95.3 | 52.0 | 43.3% |
| Turns + Jev, 3 | 66.8% | 534 | 77.8 | 47.9 | 98.2 | 46.7 | 98.4 | 53.5 | 43.3% |
| Generous budget ( 20, and full context) | |||||||||
| Turns + cosine, 20 | 72.8% | 4,352 | 84.7 | 58.7 | 98.2 | 40.0 | 96.9 | 63.8 | 70.0% |
| Turns + Jev, 20 | 73.8% | 2,278 | 84.7 | 67.8 | 96.4 | 46.7 | 96.9 | 58.3 | 63.3% |
| System | Write $/1k turns | LLM | Jev | Embeddings | Write p50 (s) | Read $/query | Jev calls/query | Read p50 (ms) | Read p90 (ms) |
|---|---|---|---|---|---|---|---|---|---|
| Turns + Jev | $0.0006 | – | – | $0.0006 | 0.2 | $0.00020 | one | 273 | 359 |
| Turns + cosine | $0.0006 | – | – | $0.0006 | 0.2 | 0 † | none | 219 | 263 |
| Turns + LLM | $0.0006 | – | – | $0.0006 | 0.2 | $0.00023 | one (no-op) | 816 | 1,107 |
| engram v2 | $1.865 | $1.322 | $0.542 | $0.0015 | 2.0–2.1 | $0.00018 | one | 276 | 319 |
| mem0 | $1.321 | $1.319 | – | $0.0017 | 2.1–2.3 | 0 † | none | 486 | 718 |
| Jev-Mem | $0.212 | – | $0.211 | $0.0011 | 0.5 | $0.00120 | 5.3 | 1,329 | 1,622 |
| System | 3 | 20 |
|---|---|---|
| Turns + cosine | 63.6% | 47.8% |
| engram v2 | 59.8% | 52.2% |
| mem0 | 59.8% | 53.6% |
| Turns + Jev | 54.1% | 49.3% |
| Turns + LLM | 48.8% | 53.1% |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| System, setting | conv-44 (123) | conv-47 (150) | conv-48 (191) | conv-49 (156) | conv-50 (158) |
|---|---|---|---|---|---|
| Turns + Jev, 3 | 78.0 | 76.7 | 79.1 | 73.7 | 78.5 |
| Turns + Jev, 6 | 79.7 | 76.7 | 78.5 | 73.7 | 76.6 |
| Turns + Jev, 20 | 79.7 | 74.7 | 81.7 | 75.0 | 76.6 |
| Turns + cosine, 3 | 65.9 | 57.3 | 62.8 | 58.3 | 55.7 |
| Turns + cosine, 20 | 78.0 | 74.0 | 79.1 | 74.4 | 74.7 |
| Turns + LLM, 3 | 76.4 | 72.7 | 80.6 | 77.6 | 79.7 |
| Date | Change | Reason | Effect on results |
|---|---|---|---|
| 2026-09-26 | Amendment, deposited after Batch A (S1 and S2 known) and before any H1 result: shortlist recall; LongMemEval on all 500 questions with user and assistant turns, with the new test S7; the second answer model; the outcome paragraphs and the rule for “LongMemEval holds”; budget caps | Extend the study before the primary test was run | S7 joins the Holm family; new robustness checks; H1, its margin and the other tests unchanged |
| 2026-09-26 | Outcome paragraphs revised before upload | The first author’s own wording | None; made before any H1 result |
| 2026-09-26 | mem0’s extraction calls went through OpenRouter, about half served by Azure, instead of the OpenAI API; the ledger was corrected and guards added before any later run | A mem0 library default routes calls to OpenRouter when its key is set ( Appendix G ) | S3 is reported with a caveat; no other test involves mem0 |
| 2026-09-26 | Runs repeated after OpenAI rate limits, with more retries and a Jev throttle; completed calls replayed from the call cache | Rate limits | None: replayed calls are identical |
| 2026-09-26 | Turns + LLM also asks Jev’s query-relation question once per query | Shared read-path code | None: the question is a no-op on turns |
| 2026-09-26 | Token counting treats text that spells a special token ( <|endoftext|> , in one LongMemEval haystack) as ordinary text | The tokenizer refused that text | One question re-run; every other count unchanged |
| Mapping | Agreement with judge | On Turns + Jev’s answers | On engram v2’s answers | Only Turns + Jev right | Only engram v2 right | Both right | Both wrong |
| Strict | 81% | 81% | 82% | 47 | 59 | 17 | 18 |
| Lenient | 79% | 82% | 77% | 41 | 60 | 31 | 9 |
| Mapping, treatment | Questions | Turns + Jev | engram v2 | Difference | One-sided 95% bound | Two-sided 95% CI |
|---|---|---|---|---|---|---|
| Strict, (a) judge’s labels | 778 | 76.3% | 78.0% | 1.7 | 3.9 | [ 4.3, +0.9] |
| Strict, (b) dropped | 777 | 76.4% | 78.0% | 1.5 | 3.7 | [ 4.1, +1.1] |
| Lenient, (a) judge’s labels | 778 | 77.4% | 79.9% | 2.6 | 4.7 | [ 5.1, 0.03] |
| Lenient, (b) dropped | 777 | 77.5% | 79.9% | 2.4 | 4.6 | [ 5.0, +0.09] |
| Category | Questions | All evidence in shortlist | At least one | Rerank keeps none |
|---|---|---|---|---|
| Multi-hop | 250 | 42.8% | 88.4% | 5.9% |
| Temporal | 284 | 83.5% | 88.4% | 22.7% |
| Open-domain | 83 | 42.0% | 63.0% | 31.4% |
| Single-hop | 771 | 89.2% | 91.2% | 5.8% |
| All | 1,388 | 76.9% | 88.5% | 10.4% |