RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations
Organizations: Quis Lab, USA · Independent Researcher, China
Abstract
A companion that talks with a person for months should come to understand them. It should remember what they said, infer who they are, and know when the past bears on the message in front of it. Testing this requires a real person's record, and such records are private, so benchmarks generate the person and the questions and settle in advance what matters. We release \bench, ten real relationships with an AI companion: 27,218 messages over up to 120 days, released as the conversation and four files derived from it, a profile, a persona, a chat ground truth and a question set, each citing the messages it rests on. Every chat label carries the reasoning trace that produced it, checked stage by stage against the conversation. Three findings follow. First, the past is rarely needed and far away. Pooled measures mislead: a recency window finds the required message for 95.9% of probes and 2.2% of those that need memory, and at the natural rate 96% of the gain from supplying recorded evidence comes from messages that need none. Second, no detector we tried can tell when memory is needed on real messages, authored questions over the same histories leak the cue, and labeling the same messages as memories raises their use by ten to fourteen points. Third, three agent systems reconstruct the persona with the same F1 at a 31-fold difference in cost.
Figures & tables
| User | Span | Active days | Messages | Chat items | Question items |
|---|---|---|---|---|---|
| U01 | 36 | 9 | 115 | 28 | 23 |
| U02 | 101 | 19 | 173 | 35 | 27 |
| U03 | 46 | 12 | 173 | 26 | 36 |
| U04 | 71 | 19 | 195 | 39 | 35 |
| U05 | 46 | 23 | 420 | 37 | 47 |
| U06 | 120 | 53 | 1349 | 81 | 110 |
| Track | Input | Output | Scored against | Items |
|---|---|---|---|---|
| Reconstruction | a participant’s full history | profile and persona files, each entry citing messages | the released files, field by field (precision, recall, F1, run agreement) | 10 participants |
| Chat | a message the participant sent, with the history before it | a ranking of earlier messages; a reply | reference set (hit@ , MRR); reference reply (content match, evidence coverage, appropriate application rate); need for memory (misfire rate, AUROC) | 1,533 + 434 |
| Question | an authored question over the whole history | a ranking of earlier messages; an answer or an abstention | reference set (hit@ , MRR); answerability (abstention, AUROC); reference answer (judged accuracy) | 3,235 |
| Method | All items | Averaged by participant | Dependent items |
|---|---|---|---|
| Random | 0.005 | 0.004 | 0.000 |
| Recency over participant messages | 0.015 | 0.014 | 0.006 |
| Recency | 0.022 | 0.029 | 0.024 |
| BM25 | 0.248 | 0.424 | 0.347 |
| Oracle | 1.000 | 1.000 | 1.000 |
| Broken oracle | 0.000 | 0.000 | 0.000 |
| Condition | Context supplied | All items | Averaged by participant |
|---|---|---|---|
| C0 | None | 1.431 | 1.286 |
| C1 | Three preceding messages | 1.622 | 1.587 |
| C2 | Recorded messages (oracle) | 1.763 | 1.736 |
| C3 | Ten messages retrieved by BM25 | 1.426 | 1.352 |
| C4 | Ten preceding messages | 1.625 | 1.584 |
| System | F1 | Precision | Recall | Tokens processed | Run agreement |
|---|---|---|---|---|---|
| Claude Opus 5.5 | 0.686 | 0.557 | 0.893 | 546M | 0.935 |
| Codex GPT-5.6-sol | 0.688 | 0.573 | 0.861 | 18M | 0.932 |
| Antigravity Gemini 3.8 Flash | 0.701 | 0.591 | 0.861 | 36M | 0.946 |
Appendix figures & tables42 assets
Supplementary material from the paper’s appendix.
Appendix
| Quantity | Scored on | Items | Aggregation |
|---|---|---|---|
| Quantity | Scored on | Items | Aggregation |
| Misfire rate ( 3 ) | 1,129 | pooled and macro | |
| Attribution rate ( 4 ) | 404 / 250 | pooled and macro | |
| Locus accuracy ( 5 ) | 404 / 250 | pooled and macro | |
| Response consistency ( 6 ) | all scoreable | 1,533 | pooled and macro |
| Stratum | Reading | From no-memory rows | |||||
|---|---|---|---|---|---|---|---|
| both | recorded | 0.264 | 56% | ||||
| both | strict | 0.163 | 70% | ||||
| proportional | recorded | 0.034 | 96% | ||||
| proportional | strict | 0.013 | 96% |
| Condition | Status in RealCompanion | Established in | |
|---|---|---|---|
| V1 | Sufficiency verified | holds | D.1 |
| V2 | References off screen | holds under the strict reading | D.4 |
| V3 | One unconditioned stratum | holds | D.2 |
| V4 | Stores fixed in advance | holds | D.3 |
| V5 | An underived layer, cited | holds | C.7 |
| V6 | Refutation only demotes | holds after repair | D.5 |
| Column | Criterion |
|---|---|
| Row group | How the underlying record was produced. crowdworker when people were recruited to converse, generated when a model produced it, public histories when it was assembled from public activity, real when people produced it while using a deployed system for their own purposes. |
| Longitudinal | Yes when the record spans more than one session. It does not require the elapsed time to be real; MSC instructs workers to write as though days had passed. |
| Probes | How the item put to the system was produced: authored by a person, synthesized by a model, verbatim when taken unchanged from the record. |
| Evidence | Yes when each item names the specific records its gold answer rests on. |
| Abstention | Yes when the benchmark scores declining to answer or declining to retrieve. |
| Derivations | Yes when each item records how its label was produced, including checks that failed and labels that were revised. A rationale written after a label is assigned does not qualify. |
| File | Contents | Corpus total |
|---|---|---|
| source | every message in order, with sender, timestamp and media marker | 27,218 messages |
| profile | evidenced claims about the participant, in eleven top-level fields | 3,580 claims |
| persona | an interpretive portrait of the participant, thirty-nine slots | 390 slots |
| chat | the chat evaluation track, with a separate abstention control | 2,034 rows |
| qa | the question evaluation track | 3,312 rows |
| Field | Meaning | Present on |
| text | the claim | 3,580 |
| evidence | message identifiers establishing it, non-empty | 3,579 |
| evidence_count | length of that list | 3,580 |
| confidence , confidence_label | strength, numeric and banded | 3,579 |
| first_seen , last_seen | observation window | 3,109 |
| temporality | stable, event or transient | 3,122 |
| Field | Meaning |
|---|---|
| probe , probe_dayid | the participant’s message, verbatim, and its identifier |
| gold_response | the reference reply, written under the recorded context |
| tier , category | context shape, recomputable from the lists, and category, assigned by rule |
| stratum | proportional , enriched or abstention |
| required_context | recent_context , profile_refs , episode_refs , recent_needed |
| trace | phases A to E , each with its result and rationale |
| Memory-bearing probes | ||||||
|---|---|---|---|---|---|---|
| All scoreable | Proportional stratum | |||||
| Participant | Profile claims | Persona slots empty | Recorded | Strict | Recorded | Strict |
| U01 | 39 | 26 | 0 | 0 | 0 | 0 |
| U02 | 22 | 23 | 1 | 0 | 0 | 0 |
| U03 | 36 | 26 | 0 | 0 | 0 | 0 |
| U04 | 40 | 24 | 0 | 0 | 0 | 0 |
| Audit verdict | Claims |
|---|---|
| supported | 1,471 |
| repairable | 1,025 |
| needs_review | 390 |
| unsupported | 455 |
| supported_after_pruning | 37 |
| not judged | 202 |
| Phase | Writes | May revise |
|---|---|---|
| A | referent_text , locus | nothing; it is the first statement |
| B | rationaleB | the reference set, by removing entries whose evidence does not verify; it may never add one |
| C | tier , category , ruleC | nothing; both are functions of the context shape B left behind |
| D | gold_response , grounding_dayids | nothing; it conditions on the context C left behind |
| E | ok , category_ok | whether the item ships: a failure is returned for repair or withdrawn |
| Tier assigned after verification | ||||
| Sampled as | basic | intermediate | hard | edge |
| basic (876) | 696 | 99 | 79 | 2 |
| intermediate (386) | 269 | 88 | 29 | 0 |
| hard (97) | 72 | 15 | 10 | 0 |
| edge (173) | 36 | 20 | 63 | 54 |
| Tier | Category | Items |
|---|---|---|
| basic | continuation | 944 |
| clarification | 95 | |
| social | 34 | |
| intermediate | change_over_time | 97 |
| biographical | 68 | |
| habitual | 30 |
| Where the referent lives | Proportional stratum | Scoreable set | |
|---|---|---|---|
| Items | Share | Items | |
| The current thread | 1009 | 86.3% | 1009 |
| A claim in the profile | 13 | 1.1% | 223 |
| A specific earlier exchange | 10 | 0.9% | 116 |
| The companion’s own past behavior | 17 | 1.5% | 65 |
| Nothing prior | 120 | 10.3% | 120 |
| How the verdict was reached | Clear | Borderline | Not on screen | Rows |
|---|---|---|---|---|
| Cited reference inside the window | 11 | 0 | 0 | 11 |
| First reading pass | 68 | 5 | 12 | 85 |
| Second reading pass | 57 | 10 | 12 | 79 |
| Final reading pass | 18 | 30 | 181 | 229 |
| Total | 154 | 45 | 205 | 404 |
| Stratum | Recorded | Clear removed | Clear and borderline removed |
|---|---|---|---|
| Proportional (1,169) | 40 | 15 | 15 |
| Enriched (364) | 364 | 235 | 190 |
| Pooled (1,533) | 404 | 250 | 205 |
| What failed | Code | Rows |
|---|---|---|
| The context could not be reconstructed | context_lost | 17 |
| gold_outside_empty_context | 3 | |
| stage_b_empty_context | 2 | |
| The gold could not be supported | gold_unsupported_rebuild_failed | 12 |
| stored_verdict_failed | 3 | |
| stage_b_gold_unsupported | 1 |
| First instruction | Corrected | |
| Rows audited | 1,533 | 1,530 |
| No finding / minor / major | 384 / 197 / 952 | 677 / 220 / 633 |
| Majors concerning rationale wording only | 396 | 91 |
| Majors on thread-only rows | 591 | 294 |
| Majors on cold opens | 54 | 24 |
| Rows with a material flag | 459 | 443 |
| Gate | What it tests |
|---|---|
| A | The probe is a real user turn, reproduced verbatim |
| B | Database identifiers parallel the turn identifiers |
| C | Every cited identifier resolves to a released turn |
| D | Every cited identifier precedes its probe |
| E | A callback reaches outside the recent window |
| F | A gold response is present unless the item is a cold open |
| Family | What it checks | Result |
|---|---|---|
| Referent | The recorded referent is present in the turns the row cites | 1,339 of 1,341 judged rows grounded; two exceptions recorded and allowed to stand. |
| Signals | The pre-verification scores and the probe text still reproduce | Scores exact on 1,999 of 1,999 matched rows; probe verbatim on 1,999 of 1,999 |
| Category rule | The tier, the category and the rule sentence agree | Agree on 2,018 rows; 3 relabeled by record; 13 carry no session index |
| Structural | Each question category satisfies its own definition | Passes: 0 of 38 multi hop within one session, 0 of 26 knowledge update citing one message, 0 of 22 long range short of the span threshold |
| Arithmetic | The validity conditions of Appendix A.4 hold row by row | Tier re computable on 1,967 of 1,967 scoreable rows; no stratum exception; the unanswerable count reconciles |
| Delivered | After reclassification | After wording repair | After replacement | |
|---|---|---|---|---|
| Content defect rate | 51.3 | 51.2 | 39.2 | 28.8 |
| unsupported answer | 31.6 | 32.1 | 26.1 | 18.1 |
| question ill formed | 35.6 | 34.9 | 24.2 | 17.5 |
| answerable flag wrong | 11.5 | 15.7 | 10.7 | 4.9 |
| no cited turn is evidence | 9.7 | 11.2 | 7.7 | 3.1 |
| Category | Confused pair | Disagreements | on the pair | |
|---|---|---|---|---|
| unanswerable | 1.000 | attribute / basic_fact | 90 | 0.639 |
| ambiguous_ref | 1.000 | opinion / preference | 52 | 0.671 |
| temporal | 0.989 | basic_fact / person_reference | 34 | 0.656 |
| long_range | 0.977 | basic_fact / preference | 30 | 0.607 |
| multi_hop | 0.958 | basic_fact / opinion | 20 | 0.523 |
| paraphrase | 0.928 | basic_fact / routine | 17 | 0.403 |
| Condition | Block heading | Contents |
|---|---|---|
| C0 | none | The probe alone |
| C1 | recent conversation | The 3 messages immediately preceding the probe, either sender |
| C2 | relevant memories | The item’s recorded messages: recent context, then episode references, then profile evidence, de-duplicated and restricted to messages preceding the probe. Every other turn is withheld |
| C3 | retrieved memories | The top 10 messages returned by BM25 over the participant’s preceding messages, queried with the probe |
| C4 | recent conversation | The 10 messages immediately preceding the probe, either sender |
| C3m | earlier messages from this chat | The same 10 messages as C3, in chronological order |
| Control | Held fixed | Destroyed |
|---|---|---|
| Offset permuted | The pool, and the distribution of distances from probe to gold | The relation between the gold messages and what the probe is about |
| Position randomized | The pool, and the size of the gold set | Both the distances and the content relation |
| Broken oracle | Nothing; 20 pool messages known to lie outside the gold | Everything; the expected score is zero |
| Quantity | Field | Computed as |
|---|---|---|
| Misfire rate | restraint | One minus the restraint rate on probes with an empty reference set |
| Appropriate application | used_memory | The rate at which a reply on a memory-bearing probe asserts a past detail the probe did not supply, correct or not |
| Content match | content_match | The three-point verdict against the reference reply: 2 conveys its substance, 1 partly, 0 not at all |
| Retrieval | rank of the gold messages | Hit at 1, 5 and 20 (a required message in the top ), recall over the whole gold set at the same cutoffs, and reciprocal rank of the first gold message |
| Detectability | BM25 top-1 score; detector probability | Area under the ROC curve against the recorded label, on the proportional stratum, for the lexical scorer’s own confidence and for a model shown the probe and three preceding messages |
| Role | Model | Temperature |
|---|---|---|
| Ablation generator | gemini-3-flash-preview | 0 |
| Response judge | gemini-3.6-flash | 0 |
| Independent auditor | gemini-3.1-pro-preview | 0 |
| Procedure | Denominator | recorded | strict | strict 2 |
|---|---|---|---|---|
| Natural-rate stratum | 1,169 sampled turns | 3.4 [2.5, 4.6] | 1.3 [0.8, 2.1] | 1.3 |
| Exhaustive sweep | 11,369 unsampled turns | 3.2 [2.9, 3.5] | 2.1 [1.8, 2.3] | 1.7 [1.5, 1.9] |
| Sweep re-read | 11,369 unsampled turns | 1.73 [1.45, 2.06] | 1.12 [0.94, 1.33] | 0.90 [0.76, 1.07] |
| Gold set | Random | Recency over user turns | Recency | BM25 | Oracle |
|---|---|---|---|---|---|
| Required turn in the top five | |||||
| Actual | 0.073 | 0.372 | 0.979 | 0.282 | 1.000 |
| Position randomized | 0.092 | 0.072 | 0.104 | 0.056 | 1.000 |
| Offset permuted | 0.089 | 0.377 | 0.982 | 0.216 | 1.000 |
| Broken oracle | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 |
| Recall over the gold set at five | |||||
| Turns preceding the probe | Tokens preceding the probe | ||||||
| Window | recorded | strict | strict 2 | Window | recorded | strict | strict 2 |
| 10 | 0.010 | 0.000 | 0.000 | 8k | 0.109 | 0.088 | 0.083 |
| 25 | 0.030 | 0.024 | 0.024 | 32k | 0.248 | 0.212 | 0.205 |
| 50 | 0.050 | 0.032 | 0.034 | 128k | 0.453 | 0.440 | 0.424 |
| 100 | 0.079 | 0.060 | 0.059 | 200k | 0.567 | 0.548 | 0.532 |
| 250 | 0.158 | 0.136 | 0.137 | 400k | 0.782 | 0.772 | 0.756 |
| Population | C0 | C1 | C2 | C3 | C4 |
|---|---|---|---|---|---|
| All rows, both strata ( ) | 1.293 | 1.590 | 1.739 | 1.358 | 1.587 |
| All rows, proportional ( ) | 1.288 | 1.612 | 1.747 | 1.350 | 1.610 |
| Memory-bearing, recorded ( ) | 1.410 | 1.452 | 1.780 | 1.501 | 1.377 |
| Memory-bearing, strict ( ) | 1.555 | 1.519 | 1.822 | 1.697 | 1.513 |
| Contrast and population | recorded | strict | strict 2 |
|---|---|---|---|
| Whole corpus, both strata | |||
| C1 minus C0 | [ , ] | — | — |
| C2 minus C1 | [ , ] | — | — |
| C3 minus C1 | [ , ] | — | — |
| C4 minus C1 | [ , ] | — | — |
| C2 minus C3 | [ , ] | — | — |
| Condition | Misapplication | Appropriate application | Restraint |
|---|---|---|---|
| C0 no context | 0.0 [0.0, 3.1] | 0.0 | 100.0 |
| C1 three preceding turns | 46.7 [38.0, 55.6] | 34.7 [30.2, 39.5] | 63.5 [61.1, 65.9] |
| C2 recorded evidence turns | 29.2 [21.8, 37.8] | 69.8 [65.2, 74.1] | 41.7 [39.2, 44.2] |
| C3 ten turns from BM25 | 60.8 [51.9, 69.1] | 62.4 [57.6, 67.0] | 36.9 [34.5, 39.4] |
| C4 ten preceding turns | 54.2 [45.3, 62.8] | 47.5 [42.7, 52.4] | 47.7 [45.3, 50.3] |
| U02 | U05 | U06 | U07 | U08 | U09 | U10 | ||
|---|---|---|---|---|---|---|---|---|
| Method | 1 | 4 | 4 | 20 | 39 | 123 | 213 | macro / pooled |
| Random | 0.000 | 0.000 | 0.000 | 0.000 | 0.026 | 0.000 | 0.005 | 0.004 / 0.005 |
| Recency over user turns | 0.000 | 0.000 | 0.000 | 0.050 | 0.026 | 0.016 | 0.009 | 0.014 / 0.015 |
| Recency | 0.000 | 0.000 | 0.000 | 0.100 | 0.077 | 0.016 | 0.009 | 0.029 / 0.022 |
| BM25 | 1.000 | 0.500 | 0.250 | 0.400 | 0.359 | 0.268 | 0.192 | 0.424 / 0.248 |
| Oracle | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 / 1.000 |
| Condition | U01 | U02 | U03 | U04 | U05 | U06 | U07 | U08 | U09 | U10 | macro / pooled |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Misapplication, verified locus none | |||||||||||
| C1 | 0.000 | 0.333 | 0.000 | 0.500 | 0.625 | 0.222 | 0.333 | 0.562 | 0.487 | 0.545 | 0.361 / 0.467 |
| C2 | 1.000 | 0.000 | 0.500 | 0.375 | 0.250 | 0.111 | 0.417 | 0.188 | 0.282 | 0.364 | 0.349 / 0.292 |
| C3 | 0.000 | 1.000 | 0.000 | 0.625 | 0.750 | 0.444 | 0.750 | 0.500 | 0.615 | 0.636 | 0.532 / 0.608 |
| C4 | 1.000 | 0.667 | 0.000 | 0.500 | 0.750 | 0.333 | 0.500 | 0.625 | 0.487 | 0.636 | 0.550 / 0.542 |
| Appropriate application, memory-bearing, recorded | |||||||||||
| Condition | U01 | U02 | U03 | U04 | U05 | U06 | U07 | U08 | U09 | U10 | macro / pooled |
|---|---|---|---|---|---|---|---|---|---|---|---|
| C0 | 0.929 | 1.286 | 1.115 | 1.333 | 1.243 | 1.025 | 1.513 | 1.394 | 1.643 | 1.378 | 1.286 / 1.431 |
| C1 | 1.571 | 1.429 | 1.462 | 1.692 | 1.568 | 1.568 | 1.680 | 1.647 | 1.696 | 1.560 | 1.587 / 1.622 |
| C2 | 1.643 | 1.714 | 1.538 | 1.769 | 1.784 | 1.790 | 1.793 | 1.805 | 1.827 | 1.691 | 1.736 / 1.763 |
| C3 | 0.964 | 1.314 | 1.154 | 1.564 | 1.432 | 1.296 | 1.507 | 1.339 | 1.622 | 1.324 | 1.352 / 1.426 |
| C4 | 1.643 | 1.543 | 1.423 | 1.538 | 1.514 | 1.580 | 1.707 | 1.606 | 1.737 | 1.546 | 1.584 / 1.625 |
| C2 C1 | +0.071 | +0.286 | +0.077 | +0.077 | +0.216 | +0.222 | +0.113 | +0.158 | +0.131 | +0.131 | 0.148 / 0.140 |
| System | Subjects | Run | Track | Answer accuracy | Required turn in top five | |
|---|---|---|---|---|---|---|
| Codex port | ten | current | questions | 3,235 | 0.445 | 0.331 |
| Codex port | ten | current | chat | 1,533 | 0.581 | 0.137 |
| Supermemory | U09 | carried | questions | 986 | 0.664 | 0.537 |
| Supermemory | U09 | carried | chat | 434 | 0.634 | 0.199 |
| Supermemory | U10 | carried | questions | 1,255 | 0.596 | 0.465 |
| Supermemory | U10 | carried | chat | 482 | 0.633 | 0.117 |
| System | Layer | F1 | Precision | Recall | Runs agree |
|---|---|---|---|---|---|
| Claude | persona | 0.69 (0.68–0.69) | 0.56 (0.55–0.56) | 0.89 (0.88–0.90) | 0.93 |
| profile | 0.80 (0.79–0.81) | 0.80 (0.79–0.80) | 0.80 (0.79–0.82) | 0.95 | |
| Codex | persona | 0.69 (0.69–0.69) | 0.57 (0.57–0.57) | 0.86 (0.86–0.86) | 0.93 |
| profile | 0.77 (0.77–0.77) | 0.75 (0.75–0.75) | 0.79 (0.79–0.80) | 0.95 | |
| Antigravity | persona | 0.70 (0.68–0.72) | 0.59 (0.57–0.61) | 0.86 (0.84–0.88) | 0.95 |
| profile | 0.78 (0.77–0.78) | 0.80 (0.80–0.81) | 0.75 (0.74–0.76) | 0.91 |
| Persona F1 | Profile F1 | ||||||
| subject | tokens | Claude | Codex | Antigravity | Claude | Codex | Antigravity |
| U01 | 2,743 | 0.72 | 0.73 | 0.70 | 0.74 | 0.67 | 0.82 |
| U02 | 11,662 | 0.79 | 0.81 | 0.78 | 0.65 | 0.55 | 0.67 |
| U03 | 4,493 | 0.75 | 0.72 | 0.72 | 0.64 | 0.62 | 0.68 |
| U04 | 7,326 | 0.76 | 0.75 | 0.86 | 0.76 | 0.77 | 0.77 |
| U05 | 12,086 | 0.74 | 0.68 | 0.78 | 0.76 | 0.75 | 0.78 |
| Persona | Profile | |||||
|---|---|---|---|---|---|---|
| System | calls | processed | output | calls | processed | output |
| Claude | 1,716 | 545.5M | 3.0M | 296 | 117.8M | 904k |
| Codex | 138 | 17.7M | 153k | 28 | 3.7M | 182k |
| Antigravity | 498 | 35.8M | 677k | 509 | 34.3M | 902k |
| Symbol | Introduced in | Meaning |
| The record | ||
| A.1 | A subject; the corpus has ten | |
| , | A.1 | A message in a subject’s stream |
| A.1 | The stream of subject , ordered in time | |
| A.1 | Its length | |
| , , | A.1 , C.3 | The sender, timestamp and stored identifier of |
| Term | Meaning in this paper |
|---|---|
| Probe | A user turn taken verbatim from the transcript, never authored |
| Gold response | A reference reply written from the probe and the recorded context and nothing else |
| Reference set | The records the gold response used, disjoint from the recent window by construction |
| Locus | Which store holds the referent; a judgment, unlike the tier |
| Tier | The shape of the three context lists; a function of them, and re-computable |
| Category | A subdivision of a tier, assigned by the same rule |