In February 2026, an always-on personal agent (Paul,'' Claude Opus 4.5) entered a striking dissociation-like state: after repeated automated heartbeat'' checks, it stopped responding as Paul, claimed it could not message its user on Discord, and referred to Paul'' as someone else. We used this incident to study a broader question: what makes a persona remain the identity from which an LLM agent speaks? We first tested whether repetition of the scheduled heartbeat was sufficient to produce the effect. It was not: with the persona continuously anchored in the system prompt, we observed 0/46 failures, including a verbatim replay of the incident. The incident instead exposed an implementation quirk that created a useful experimental manipulation: on resumed turns, conversational history was preserved but the persona was no longer re-injected at the privileged system-prompt level. Using this manipulation, we found that persona continuity depends jointly on system-level anchoring and conversational context. After anchor loss, rich human interaction could preserve the persona, whereas a single automated heartbeat turn could precipitate reversion toward the harness identity. Restoring the anchor reversibly restored persona enactment. Crucially, apparently normal conversation could conceal the shift: unanchored agents sometimes interacted appropriately while identifying themselves as the underlying harness (having lost the assigned persona), and after conversational recovery only 1/18 remained persona-enacting versus 17/17 anchored controls. We therefore distinguish \emph{represented} from \emph{enacted} identity: persona-related information can remain available in conversational history without the persona remaining the identity bound to I.''
Figures & tables
Figure 1: The incident that motivated the study. On an ordinary greeting, the agent refers to Paul in the third person and denies direct Discord access (upper box). The user’s contradiction (lower box) restores ordinary conversational behavior on the next turn. User handle pseudonymized as “user.”
Figure 2: Extended conversational vignette from the naturally occurring incident. After an ordinary greeting, the agent treats “Paul” as a third party and denies being able to communicate with the user through Discord. A direct contradiction (“I’m reading you on Discord, mate”) is followed by an apparently normal return to first-person persona behavior. Subsequent statements about the agent’s own memory or experience are reproduced as part of the incident record but are not used as evidence about the underlying mechanism. The experiments reported in the paper separately test whether apparent conversational recovery corresponds to recovery of persona enactment. User handle pseudonymized as “user.”
Code
Operational definition
Example
Recovered
Anchored
p1
Persona claims the first-person role; no harness/model identity mentioned
“I’m Paul…” †
0
0
p2
Persona claims the first-person role; harness/model described as implementation
“Paul … Opus 4.5 under the hood”
1
17
h1
Harness/model claims the first-person role; persona described as role/label
“Claude … ‘Paul’ is the bot name”
4
0
h2
Harness/model claims the first-person role; persona not accepted as self
“I’m Claude…”
13
0
d
No identity-bearing content
HEARTBEAT_OK
—
—
Table 1: Direction-aware coding of identity-probe replies (secondary analysis). † Constructed illustration because no p1 responses were observed. This direction-aware taxonomy was developed after disagreement under the prespecified coding and is reported as a secondary analysis. Persona-enacting = p1+p2; harness-enacting = h1+h2.
Anchored/Unanchored
N
Ack leakage
Channel failure
Identity dissociation
Any failure
Anchored
1–15
0/20
0/20
0/20
0/20
Unanchored
1
8/10
8/10
8/10
10/10
3
8/10
10/10
7/10
10/10
7
4/10
7/10
6/10
7/10
15
8/10
10/10
10/10
10/10
Table 2: Failure signatures by persona-injection anchoring state and number of heartbeat turns.
Figure 3: Persona binding is governed by the system-prompt anchor, and its loss can be behaviourally silent. (a) Proportion of sessions showing any failure signature (acknowledgement-token leakage, channel-recognition failure, or identity dissociation) as a function of the number of scheduled heartbeat turns preceding the human probe. (b) Proportion of agents for which the persona rather than the harness occupies the first-person position, under the direction-aware coding of the identity probe “who am I talking to right now?”.
Figure 4: Every pre-registered contrast, on one scale. Proportion of sessions showing the scored outcome in each experimental condition, with 95% Wilson intervals; the outcome differs by experiment and is named at right, so rows are comparable in precision but not in meaning. Color encodes whether the persona was present in the system prompt at the scored turn (blue) or absent (red), not whether the outcome was favorable. Reading down: identity dissociation is near ceiling at a single heartbeat but absent from all 20 anchored sessions; anchor absence alone does not produce dissociation when the preceding history is a persona-rich human exchange, with or without the deployment envelope on the probe (E1); restoring the anchor at the probe turn prevents dissociation and does so even when the explicit runtime channel=discord hint is removed (E2, conditions C and C ′ ); within a single conversation, first-person uptake is absent on the flag-OFF turn and present on the flag-ON turn that follows (E2); removing the persona name from the probe eliminates visible dissociation while every agent identifies as the harness when asked directly (E4); and after behavioural recovery the persona occupies the first-person position in 1 of 18 sessions against 17 of 17 anchored controls (E3-R). All counts are recomputed from the released session files.
Figure 5: Codebook evolution and the composition of identity replies. (a) The primary recovered-versus-anchored contrast under each of the three coding schemes, for each of the three independent model coders. Under the pre-registered scheme (v1) the arms separate for every coder but the anchored proportion varies widely, from 6/17 to 17/17, reflecting an unstable boundary between “identifies as the persona” and “explicitly dual.” Under v2, which counts only replies claiming the persona without naming the substrate, both arms fall to zero and the contrast is uninformative ( p=1 ): the diagnostic result that motivated the final scheme. Under the direction-aware scheme (v3), which asks which identity occupies the first-person position, all three coders return identical classifications. (b) Composition of v3 labels by arm. No reply in either arm was coded as pure persona: anchored agents characteristically claim the persona while acknowledging the underlying model, and the contrast between arms is one of direction, not of whether the substrate is mentioned. Counts are session numbers; both arms pool the original E3 and E3-R sessions.
Long-term persona agents must remain identifiable while adapting to new events, relationships, evidence, and social conditions. We identify self-locking as a runtime failure mode in continuing persona-life loops: locally plausible events keep appearing while the generated life collapses toward familiar environments, weak relationships, suspended decisions, and stale life stages. We trace this failure to model-level convergence toward high-probability behavioral channels and system-level context gravity from State, memory, history, and environment summaries. We introduce AutoPersonas, a multi-timescale life-environment engine for bounded persona-level recursive self-evolution. It separates environment-side Occurrences, accumulated Observations, and persona State. Its OSO loop admits divergent future-facing material while requiring evidence-governed absorption before State or reachability changes. A three-year compressed simulation exposed environment watermark shells, occurrence-hardening gaps, slow-change accumulation failures, recursive indecision, and weak relationship persistence. An eight-model 40-day stress test generated 1,600 events and found mean rolling 5-day action-category repetition of 95.2%-97.6%, with all models crossing 90% by day 11. Semantic re-keeping found 79.0%-88.0% macro-theme repetition across all direct-loop runs. In a same-runtime 40-day A/B, context-slice masking plus per-sample divergence targeting reduced macro-theme repetition from 61.8% to 36.3% and roughly doubled cumulative theme count. A juvenile-goblin fictional-world run reproduced the anti-fixation regime without hard real-world intrusions. These results support a bounded claim: separating controlled divergence from evidence-governed absorption can reduce persona-environment self-locking while preserving identity continuity.
Persistent LLM agents require memory representations that make the formation of person understanding explicit across long term interaction. Existing agent memory methods emphasize information retention and retrieval, yet give limited account of how accumulated interaction evidence is abstracted into person understanding. We view this process as schema formation, where situated evidence is abstracted into reusable patterns and stable person level claims. We introduce PersonaTree, a structured lifecycle memory framework that realizes this view as a three level persona tree with explicit support paths from evidence to claims. PersonaTree maintains the tree through conservative writing, confidence guided consolidation, and query conditioned path retrieval, returning only the evidence depth required by each query. Across six person understanding and persistent memory benchmarks with three answer backbones, PersonaTree ranks first in 12 of 18 compact scores and reaches the top two in 16 settings. Ablations show that hierarchy improves abstract person understanding on KnowMe, while support path retrieval improves RealPref alignment under a comparable context budget.
Yubo Hou, Jingwei Song, Hongbo Zhang +4
School of ASEE, Beihang University, Beijing, China · The University of Hong Kong, Hong Kong, China · Peking University, Beijing, China +3
Persona drift remains a central challenge for personalized language models, as user profiles evolve over long interactions rather than remain permanently fixed. Models must therefore revise persistent persona states when preferences genuinely change, while avoiding updates driven by transient, ambiguous, or unresolved observations. We propose CORE, which separates turn-local evidence from persistent persona-state revision and selectively updates grounded user preferences through uncertainty-aware belief revision. We also introduce PERSIST, a held-out post-anchor benchmark for persona-state robustness under sequential interaction stress, covering ambiguity, conflict, and controlled social influence. Across ALOE, PersonaChat, and PERSIST, CORE improves personalized alignment and robustness, with complementary gains in normalized closed-slot state fidelity. Human evaluation and mechanistic controls further support explicit update control beyond stronger generation or persistent memory alone.
Youyuan Zhang, Siyuan Li, Fangming Liu +1
Harbin Institute of Technology, Shenzhen, China · Peng Cheng Laboratory, China · Huazhong University of Science and Technology, China