PersMem: Internalizing Personality into Dual-Pathway Memory for LLM Agents
Authors: Hanzhong Zhang, Ziwei Xiang, Weicheng Xie, Shizhe Liu, Siyang Song
Organizations: Department of Computer Science, University of Exeter, United Kingdom · School of Marxism, Nanjing University, China · Shenzhen University, China · Department of Computer Science, University of Oxford, United Kingdom
The profile of a role-playing agent usually depends on the pre-defined personality in a system prompt, whereas its memory processing pipeline, including prioritisation of stored memories and subsequent retrieval, remains independent of this personality. This separation causes the agent's memory processing to be inconsistent with the pre-defined personality, and makes it difficult to validate whether agent behaviours follow this personality. In this paper, we propose Personality-Integrated Memory (PersMem), which integrates personality into the agent's memory processing pipeline, making it consistently personality-dependent. PersMem processes memory using four steps, where the personality is mapped to operation-specific parameters controlling: (i) affective appraisal annotating emotion states of the user input; (ii) retention of previously stored memories along with the current input; (iii) passive affect-driven memory retrieval exploring memories similar to user input in semantics and personality-guided emotions; and (iv) active goal-driven memory retrieval that refines and selects passively retrieved memories for the reply. Consequently, consistency with the pre-defined personality can be examined by inspecting memory-processing traces during human-agent interactions. We evaluate these personality-dependent differences in attachment and Big Five settings. PersMem exceeds the chance baseline for four-way attachment classification by 23.1 percentage points. In Big Five dialogue comparisons, PersMem achieves 67.5% accuracy, 6.7 percentage points above a baseline using uniformly sampled memories. On CoSER, PersMem achieves an average score of 66.13, with scores of 69.33 for Character Fidelity and 84.33 for Storyline Quality. Together, these results show that PersMem produces distinguishable personality-related memory-processing patterns.
Figures & tables
Figure 1: Design Principle and Architectural Comparison of PersMem. (a) A simplified view of PersMem. In the attachment instantiation, a fixed two-dimensional trait vector parameterises affective appraisal, memory retention, passive affect-driven memory retrieval, and the redirection of active goal-driven memory retrieval, while remaining hidden from the active controller and response model. (b) A common design specifies the assigned profile through a pre-defined personality in the system prompt while retrieving memories through a separate process. (c) PersMem instead uses one personality representation to control multiple memory operations before response generation. Fig. 2 presents the complete processing and control flow.
Figure 2: Pipeline of PersMem. (a) PersMem uses the pre-defined personality to annotate the affective state of the current user input, updates the agent’s current state, and stores the completed exchange after reply generation. (b) Memory retention controls the accessibility of stored memories, while passive affect-driven memory retrieval selects an initial memory set and proposes one candidate during each active retrieval step. (c) Active goal-driven memory retrieval either adopts the candidate proposed by passive retrieval or continues semantic retrieval and query refinement. The initial passive output and active recall set are finally merged and deduplicated.
Condition
Accuracy
95% CI / p -value
Held-out classification
Full PersMem
48.1%
[40.0%, 56.3%]; p=0.0002
Neutral gate
40.0%
–
Gate off
25.0%
–
Text embedding only
25.0%
–
Linear trait-aware reranker
30.0%
p=0.055
Table 1: Attachment profile classification. Held-out evaluation excludes internal personality-dependent scores; complete-trace diagnostics include resonance and retention scores. Primary and matched complete-trace results use different splits.
Analysis
Estimate
CI or statistic
p -value
Affective-appraisal effect
+0.005
–
–
Memory-retention effect
+0.013
–
–
Passive-retrieval effect
+0.130
–
–
Active-retrieval redirection effect
+0.085
–
–
Primary VAD classification
48.1%
[40.0%, 56.3%]
0.0002
NRC-VAD classification
45.0%
[38.8%, 51.2%]
0.0002
Table 2: Component effects and annotation checks.
Measure
Sec.
Anx.
Avo.
Fear.
Neutral OGM
–
–
0.28
0.47
Distress OGM
–
–
0.37
0.73
OGM lift over no-trait
–
–
+0.07
+0.27
Independent-label OGM difference from secure, distress
0
+0.027
+0.039
+0.067
Neutral mean recalled valence
–
–
−0.04
–
Negative-redirection proportion
0.59
0.63
0.51
–
Table 3: Profile-conditioned retrieval and response outcomes. Dashes indicate profiles excluded from an analysis.
Trait
Outcome
Full system
No-trait control
Neuroticism
Valence
−0.703
–
Extraversion
Valence
+0.776
–
Agreeableness
Valence
+0.607
–
Conscientiousness
OGM
−0.710
−0.180
Table 4: Descriptive Big Five recall associations over 50 profiles. Dashes indicate undefined valence correlations because the no-trait control produces zero variance in recalled valence.
Condition
Correct calls
Accuracy
Full system
81/120
67.5%
Random-memory control
73/120
60.8%
No-trait control
65/120
54.2%
No-memory control
64/120
53.3%
Table 5: Blind Big Five dialogue comparison. Each of 60 high–low pairs is evaluated in two presentation orders.
Method
SC
AN
CF
SQ
Avg.
CoSER Wang et al. (2025)
GPT-4o
61.59
48.93
48.95
80.33
59.95
CoSER-70B + Conv.
64.59
53.79
54.86
77.28
62.63
CogDual Liu et al. (2025)
LLaMA3.1-8B + CogDual-RL
60.10
45.89
48.82
73.08
56.97
HER Du et al. (2026)
Table 6: Role-playing results on CoSER. Published reference scores retain their source studies’ evaluation settings; our PersMem results use DeepSeek. SC, AN, CF, and SQ denote Storyline Consistency, Anthropomorphism, Character Fidelity, and Storyline Quality, respectively. Conv. denotes conversation retrieval.
Evaluation
Condition
Result
Calibration error
Fitted gain map
0.084
Fixed-gain reference
0.482
Adversarial retention
PersMem
89%
Prompt-persona baseline
58%
Table 7: Calibration and adversarial-framing results.
Method
Evidence recall@5
F1 given evidence hit
PersMem
0.387
0.547
Generative-Agents-style scoring
0.341
0.540
Semantic RAG
0.444
0.542
Table 8: LoCoMo retrieval and answer results.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Role
Value
κv
valence adjustment
0.30
κa
negative-event arousal increase
0.20
κdamp
avoidance-related arousal damping
0.15
κp
dominance increase
0.10
κq
dominance decrease
0.10
r0
base memory decay rate
10−4s−1
Appendix
Table 9: Fixed coefficients used by PersMem. Calibrated retrieval gains and the Big Five sampling temperature are defined separately in Appendices C and D.
Trait
Memory operation
Implemented relationship
N
Affective appraisal and memory retention
Stronger negative appraisal and slower forgetting of negative memories
N
Passive affect-driven memory retrieval
Stronger negative-memory term and greater weight on the complete affective relation score
N
Active-retrieval redirection
Lower gate-firing threshold
E and A
Affective appraisal and passive retrieval
Stronger positive appraisal and a stronger positive-memory term
Low A
Affective appraisal and passive retrieval
Higher dominance appraisal and a stronger preference for high-dominance memories
Low C and high N
Passive affect-driven memory retrieval
Stronger overgeneral-memory term
Appendix
Table 10: Big Five relationships used in the implemented memory operations and calibration targets.
Method
Accuracy
ΔPCB
Semantic RAG
28.3%
0
Generative-Agents-style scoring
25.8%
0
No-trait control
26.7%
0
PersMem
72.5%
+1.087
Appendix
Table 11: Complete-trace comparison on the matched analysis split. The first three methods are negative controls without personality-dependent retrieval.
Condition
Accuracy
95% CI
p -value
Full A1R1P1G1, personality gate
48.1%
[40.0%, 56.3%]
0.0002
Neutral gate
40.0%
–
–
Gate off
25.0%
–
–
Text embedding only
25.0%
–
–
Linear trait-aware reranker
30.0%
–
0.055
Appendix
Table 12: Held-out attachment profile classification. The classifier is trained on development data and fixed before test evaluation. Personality parameters and internal personality-dependent scores are excluded.
Component
Average main effect
Appraisal (A)
+0.005
Retention (R)
+0.013
Passive affect-driven memory retrieval (P)
+0.130
Gate (G)
+0.085
Appendix
Table 13: Descriptive average main effects on held-out profile classification.
Profile
Mean recalled valence
Negative-cue congruence
OGM rate
Secure
−0.12
0.60
0.23
Anxious
−0.25
0.53
0.44
Avoidant
−0.04
0.53
0.28
Fearful
−0.23
0.63
0.47
Appendix
Table 14: Mean recalled valence, negative-cue congruence, and OGM rate across attachment profiles.
Comparison
State
Difference
p
Fearful: S on − off
Neutral
+0.0422
10−5
Fearful: S on − off
Distress
+0.0770
10−5
Avoidant: S on − off
Distress
+0.0394
10−5
Anxious: S on − off
Distress
+0.0370
10−5
Anxious − secure
Distress
+0.0272
≤4×10−5
Avoidant − secure
Distress
+0.0393
≤4×10−5
Appendix
Table 15: Cross-label specificity differences computed using the second blind model-generated label set.
O
C
E
A
N
Mean
0.725
0.592
0.490
0.693
0.517
Standard deviation
0.158
0.184
0.228
0.182
0.215
Appendix
Table 16: Sample statistics for the 50 normalised IPIP FFM profiles.
Response Model and Evaluator
Full System
Random-Memory Control
No-Memory Control
No-trait control
Qwen3.5-9B and Qwen3.5-9B
72.5%
55.0%
57.5%
50.8%
Qwen3.5-9B and Llama-3.1-8B
67.5%
60.8%
53.3%
54.2%
Llama-3.1-8B and Llama-3.1-8B
65.8%
56.7%
54.2%
52.5%
Appendix
Table 17: Exploratory forced-choice accuracy across response model and evaluator configurations. Each cell contains 120 order-balanced calls over 60 underlying high–low profile pairs.
Method
Evidence recall@5
F1 given evidence hit
Semantic RAG
0.444
0.542
Generative-Agents-style scoring
0.341
0.540
PersMem
0.387
0.547
Appendix
Table 18: LoCoMo retrieval and answer decomposition. Evidence recall@5 measures retrieval coverage, while hit-conditional F1 is computed within each method’s own evidence-hit subset.
Method
SC
AN
CF
SQ
Avg.
CoSER Wang et al. (2025)
GPT-4o
61.59
48.93
48.95
80.33
59.95
Doubao-pro
60.95
49.72
47.02
79.28
59.24
Step-2
61.43
49.06
47.33
77.96
58.94
Gemini Pro
59.11
52.41
47.83
77.59
59.24
Claude-3.5-Sonnet
57.45
48.50
45.69
77.23
57.22
Appendix
Table 19: Extended comparison on CoSER. Published reference results are taken from the corresponding source studies. SC, AN, CF, and SQ denote Storyline Consistency, Anthropomorphism, Character Fidelity, and Storyline Quality, respectively.