Where the Evidence Lives: Auditing AI Companions' Self-Descriptions
Organizations: AI Consulting Division, Mamezo Co., Ltd. Tokyo, Japan · AI Technical Sector, Mamezo Co., Ltd. Tokyo, Japan · Graduate School of Artificial Intelligence and Science, Rikkyo University Tokyo, Japan
Abstract
Companion agents describe themselves: they remember, they understand their users, the relationship has changed them. We argue that such accounts, and the experience ratings that seem to confirm them, are checkable by users only where the evidence is theirs: in the agent's behavior, or in themselves. Where the evidence lives in the machinery, fluent self-description and moderately positive ratings do not establish that the mechanisms behind them ran. We demonstrate an audit procedure that sets an agent's self-description against its users' judgements and its implementation records, reporting each claim as supported, contradicted, or unresolved, and apply it to Lita, a proactive companion we built and deployed for a month with nine colleagues. Participants endorsed stylistic claims, withheld endorsement from relational ones, and rated memory at or above midpoint, while two of three memory layers had never executed their accumulation step. Memory-bearing agents should report what their self-descriptions cannot establish.
Figures & tables
if a mechanism did not run'' leads from step 3 down to optional step 4, replay and vary the criterion, and a second dashed arrow labeled diagnostic findings'' leads from step 4 up to step 5. Step 4's inputs are the fixed write stream, the repaired rule, and alternative criteria; the output is accumulation under repair and its dependence on the criterion; a note says this is a diagnostic, not a simulation of the repaired system.| Layer | Writes | Retained | Accumulation ran |
|---|---|---|---|
| Long-term (episodic facts) | 253 | 253 | yes |
| User model (behavioral patterns) | 402 | 106 | no |
| Self-narrative (identity) | 184 | 50 | no |
| Condition | Merges | Rate | Evicted | Settled |
|---|---|---|---|---|
| As deployed | 0 | 0.0% | 134 | 0 |
| LLM judge, strict | 1 | 0.5% | 133 | 0 |
| LLM judge, redundant | 47 | 25.5% | 87 | 2 |
| Reference | Pairs | Miss rate | Merge rate | Settled |
|---|---|---|---|---|
| Consensus | 58 | 4.9% | 37.5% | 5 |
| Coder A alone | 100 | 25.0% | 63.0% | 11 |
| Coder B alone | 100 | 13.3% | 52.2% | 9 |
| Evidence lies in | Claim (source) | Rating | The record shows | Outcome |
|---|---|---|---|---|
| the agent’s behavior | finds meaning in small things (NDA1) | 80.6% | not audited | unresolved: no behavioral operationalization |
| sensitive to sound (NDA2) | 72.2% | sound-related vocabulary at the participants’ rate | supported, as a behavioral tendency | |
| the machinery | remembered past talks (memory scale item) | 66.7% | 253 facts retained; 39 retrieved more than once | supported; per-respondent correspondence unchecked |
| understood my interests (memory scale item) | 50.0% | user model discarded 73.6% of observations; confidence never moved | designed accumulation contradicted; understanding by other routes unresolved | |
| the participant | changed through conversations with everyone (NDA7) | 1 of 9 agreed | self-narrative consolidation never ran | designed consolidation contradicted; change by other routes unresolved |
| understood me (relationship scale item) | 24.1% | broader item; nearest audited store is the user model | unresolved |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Criterion phrasing | Merge rate |
|---|---|
| “are these about the same underlying theme?” | 100% |
| “if the subject matter differs, answer false” | 6.7% |
| “judge the disposition drawn out, not the subject matter” (adopted) | 40% |
| Category | Items | Agree | Most | Least |
|---|---|---|---|---|
| Self-contained aesthetic | 4 | 72% | 9/9 | 1/9 |
| Relational / other-directed | 3 | 22% | 0/9 | 8/9 |
| Stratum | Msgs (M) | Retention | Reply rate | |
|---|---|---|---|---|
| High | 3 | 125.0 | 23.1% | 31.2% |
| Medium | 3 | 76.3 | 27.1% | 20.8% |
| Low | 4 | 41.0 | 38.7% | 14.9% |
| Trigger | Generated | Spoken | Rate | Answered |
|---|---|---|---|---|
| self_thought | 13,948 | 146 | 1.05% | 30 |
| conversation | 22,607 | 172 | 0.76% | 40 |
| memory_recall | 19,734 | 123 | 0.62% | 22 |
| Corpus | Named | Sensory | Sensory |
|---|---|---|---|
| /utt. | /utt. [CI] | /1k ch. | |
| Lita reactive (persona) | 1.1% | 15.6% [13.1, 18.2] | 5.35 |
| Lita proactive (persona) | 0.9% | 21.7% [17.8, 25.9] | 5.46 |
| Self-narrative (no persona) | 17.4% | 47.3% [40.2, 54.4] | 10.62 |
| Participants | 0.0% | 4.3% [2.9, 5.7] | 2.50 |