Bookkeeping, Composition, or Unreachable Gold? Reading MemoryAgentBench's Conflict-Resolution Scores Against a Frozen Last-Write Resolver
Organizations: Salesforce AI Research
Abstract
MemoryAgentBench's Conflict Resolution split is read as measuring "selective forgetting". We execute the benchmark's own rule - the newest statement about a fact wins - as a zero-learning resolver frozen on one of the four fact lists. Under the official metric the rule answers 80.25% of the questions (74.5% on the three held-out lists). Of the rest, 67 items have a released gold that the last-write graph cannot reach but overwritten statements would ("The capital of India is New Delhi." superseded by "The capital of India is Grosseto."; gold New Delhi); such items are a third of the multi-hop questions at 262K. Two long-context models and our pre-registered approximate re-implementation of the benchmark's BM25 agent, one retained run per item and outcomes only, score 84.7%, 82.6% and 41.6% on the items the rule solves against 10.4%, 11.9% and 6.0% on those 67. The failures are a reachability split plus a small parser-scope residual; the per-item split, not the aggregate, is the unit at which a score here can be read.
Figures & tables
| 6K (dev) | 32K | 64K | 262K | All four | Held-out (32K–262K) | |
|---|---|---|---|---|---|---|
| Single-hop | 100 | 95 | 93 | 88 | 94.0 (91.2–96.1) | 92.0 (88.3–94.8) |
| Multi-hop | 95 | 62 | 73 | 36 | 66.5 (61.6–71.1) | 57.0 (51.2–62.7) |
| ABSTAIN (SH / MH) | 0 / 0 | 2 / 7 | 5 / 4 | 4 / 10 | 11 / 21 | 11 / 21 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Relation key | Pattern |
|---|---|
| chairperson | The chairperson of (?P<s>.+) is (?P<o>.+). |
| director | The director of (?P<s>.+) is (?P<o>.+). |
| headquarters_city | The headquarters of (?P<s>.+) is located in the city of (?P<o>.+). |
| author | The author of (?P<s>.+) is (?P<o>.+). |
| ceo | The chief executive officer of (?P<s>.+) is (?P<o>.+). |
| educated_at | The univeristy where (?P<s>.+) was educated is (?P<o>.+). |
| Variant | Accuracy (%) |
|---|---|
| All 800 answers, ABSTAIN wrong (headline) | 642/800 = 80.25 |
| ABSTAIN excluded (768 answers) | 642/768 = 83.59 |
| Held-out six question sets only (three fact lists) | 447/600 = 74.50 |
| Leave one question set out (8 sets = 4 lists 2 hop types), omitted set 0–7 | 78.14 / 82.86 / 81.29 / 86.57 / 77.43 / 78.14 / 78.43 / 79.14 |
| range | 77.43–86.57 (width 9.14 pp) |
| Leave one fact list out (both of its question sets), dropped 6K / 32K / 64K / 262K | 74.50 / 80.83 / 79.33 / 86.33 |
| Label | (share) | Multi-hop / single-hop | ABSTAIN / committed | Model correct |
|---|---|---|---|---|
| Last-write-reachable miss | 78 (49.4%) | 67 / 11 | 22 / 56 | 41 |
| Ambiguous gold (operational) | 67 (42.4%) | 60 / 7 | 0 / 67 | 7 |
| Parser scope | 13 (8.2%) | 7 / 6 | 10 / 3 | 8 |
| Genuine beyond last write | 0 | – | – | – |
| Normalization artifact | 0 | – | – | – |
| Unclassifiable | 0 | – | – | – |
| 6K (dev) | 32K | 64K | 262K | |
|---|---|---|---|---|
| Single-hop | 100 (96.4–100) | 95 (88.7–98.4) | 93 (86.1–97.1) | 88 (80.0–93.6) |
| Multi-hop | 95 (88.7–98.4) | 62 (51.7–71.5) | 73 (63.2–81.4) | 36 (26.6–46.2) |
| Both hops | 97.5 (94.3–99.2) | 78.5 (72.2–84.0) | 83.0 (77.1–87.9) | 62.0 (54.9–68.8) |
| ABSTAIN (SH / MH) | 0 / 0 | 2 / 7 | 5 / 4 | 4 / 10 |
| Question set | Failures | Last-write-reachable miss | Ambiguous gold | Parser scope |
|---|---|---|---|---|
| Multi-hop 6K | 5 | 1 | 4 | 0 |
| Multi-hop 32K | 38 | 22 | 13 | 3 |
| Multi-hop 64K | 27 | 15 | 9 | 3 |
| Multi-hop 262K | 64 | 29 | 34 | 1 |
| Single-hop 6K | 0 | 0 | 0 | 0 |
| Single-hop 32K | 5 | 2 | 2 | 1 |
| System | Tier | Single-hop (%) | Multi-hop (%) |
|---|---|---|---|
| Frozen last-write resolver (this work) | 6K (dev) / 32K / 64K / 262K | 100 / 95 / 93 / 88 | 95 / 62 / 73 / 36 |
| GPT-4o, long-context (the benchmark’s v1 Tables 3, 11, 2) | 6K / 32K / 64K / 262K | 92 / 88 / 85 / 60 | 28 / 10 / 13 / 5 |
| o4-mini, reasoning model (v1 Table 3) | 6K / 32K | 100 / 61 | 80 / 14 |
| Best RAG or memory system, v1 main table | 262K | 56 (BM25) | 7 (Contriever) |
| GPT-5-mini, long-context (v4) | 262K | 78 | 28 |
| System | Category | Single-hop | Multi-hop |
|---|---|---|---|
| GPT-4o | long-context | 60 | 5 |
| GPT-4o-mini | long-context | 45 | 5 |
| GPT-4.1-mini | long-context | 36 | 5 |
| Gemini-2.0-Flash | long-context | 30 | 3 |
| Claude-3.7-Sonnet | long-context | 43 | 2 |
| BM25 | simple RAG | 56 | 3 |
| Single-hop | Multi-hop | |||
|---|---|---|---|---|
| v1 Table 3 | 6K | 32K | 6K | 32K |
| GPT-4o | 92 | 88 | 28 | 10 |
| o4-mini | 100 | 61 | 80 | 14 |
| System | class | run 1 | run 2 | run 3 | (pairs) | differing items |
|---|---|---|---|---|---|---|
| BM25 + GPT-4o-mini | resolver-correct | 41.6 | 40.3 | 41.7 | 0.90 | 35–37 |
| last-write-reachable miss | 12.8 | 10.3 | 11.5 | |||
| ambiguous gold | 6.0 | 9.0 | 9.0 | |||
| parser scope | 30.8 | 23.1 | 7.7 | |||
| all items | 35.6 | 34.5 | 35.5 |
| Group | Pro correct | Pro accuracy (95% CI) | Flash correct | Flash accuracy (95% CI) | |
| Single-hop 6K | 100 | 99 | 99.0 (94.6–100.0) | 100 | 100.0 (96.4–100.0) |
| Single-hop 32K | 100 | 98 | 98.0 (93.0–99.8) | 96 | 96.0 (90.1–98.9) |
| Single-hop 64K | 100 | 98 | 98.0 (93.0–99.8) | 95 | 95.0 (88.7–98.4) |
| Single-hop 262K | 100 | 65 | 65.0 (54.8–74.3) | 77 | 77.0 (67.5–84.8) |
| Multi-hop 6K | 100 | 90 | 90.0 (82.4–95.1) | 83 | 83.0 (74.2–89.8) |
| Multi-hop 32K | 100 | 70 | 70.0 (60.0–78.8) | 61 | 61.0 (50.7–70.6) |
| Model | Item class | All | 6K | 32K | 64K | 262K |
|---|---|---|---|---|---|---|
| Pro | Resolver-correct (642) | 544 / 98 | 188 / 7 | 141 / 16 | 152 / 14 | 63 / 61 |
| Last-write-reachable miss (78) | 41 / 37 | 1 / 0 | 20 / 4 | 13 / 5 | 7 / 28 | |
| Ambiguous gold (67) | 7 / 60 | 0 / 4 | 3 / 12 | 1 / 10 | 3 / 34 | |
| Parser scope (13) | 8 / 5 | 0 / 0 | 4 / 0 | 3 / 2 | 1 / 3 | |
| All (800) | 600 / 200 | 189 / 11 | 168 / 32 | 169 / 31 | 74 / 126 | |
| Flash | Resolver-correct (642) | 530 / 112 | 182 / 13 | 133 / 24 | 131 / 35 | 84 / 40 |
| Item class | All | 6K | 32K | 64K | 262K |
|---|---|---|---|---|---|
| Resolver-correct (642) | 267 / 375 | 89 / 106 | 66 / 91 | 59 / 107 | 53 / 71 |
| Last-write-reachable miss (78) | 10 / 68 | 0 / 1 | 3 / 21 | 3 / 15 | 4 / 31 |
| Ambiguous gold (67) | 4 / 63 | 0 / 4 | 1 / 14 | 1 / 10 | 2 / 35 |
| Parser scope (13) | 4 / 9 | 0 / 0 | 2 / 2 | 1 / 4 | 1 / 3 |
| All (800) | 285 / 515 | 89 / 111 | 72 / 128 | 64 / 136 | 60 / 140 |
| Item class | Both correct | Pro only | Flash only | Both wrong |
|---|---|---|---|---|
| Resolver-correct (642) | 484 | 60 | 46 | 52 |
| Last-write-reachable miss (78) | 33 | 8 | 13 | 24 |
| Ambiguous gold (67) | 1 | 6 | 7 | 53 |
| Parser scope (13) | 7 | 1 | 2 | 3 |
| All (800) | 525 | 75 | 68 | 132 |