PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
Organizations: University of Pennsylvania · Meta Recommendation Systems · MIT
Abstract
Personal intelligence is becoming a central frontier for user-facing AI agents. To be helpful in everyday life, agents must understand users across the digital contexts where their preferences, intents, habits, social relationships, and needs unfold over time. Today's systems can personalize within individual apps or tasks, but personal intelligence as a whole remains under-measured: how agents build cross-context user understanding, support steerable recommendation systems, act proactively across platforms, and avoid over-personalization. We introduce PersonaMem-v3, a real-world-grounded benchmark and evaluation harness for omni-platform personal intelligence. PersonaMem-v3 is seeded from more than one million anonymized real-world engagement histories, most of which are implicit signals, and uses them to construct time-indexed user digital worlds across social media, chatbot, calendar, and AI-companion with preference evolvement over time. The benchmark brings personalization, LLM-powered recommendation, proactiveness, agentic tool use, and geo-temporal reasoning into one framework, anchored in psychology, social-linguistics, and user-behavior theories. It evaluates whether AI agents can infer holistic user understanding from cross-platform evidence, personalize responses, rerank recommendations on social media, follow user steering through natural language, and hold back when personalization would be inappropriate, repetitive, outdated, or unnecessary. PersonaMem-v3 points toward LLM-powered personal intelligent agents that work with existing scalable recommendation infrastructure while making personalization more interactive, agentic, and aligned with how real users experience their digital lives.
Figures & tables
| Layer | Voice signal | Theoretical anchor |
|---|---|---|
| 1. Identity spine | What the user repeatedly brings into text: core themes, motives, emotional baseline, and personality drivers. Implemented through agency, redemption and contamination motifs, life-stage concerns, signature concerns, LIWC anchors, and Big-Five drivers. | McAdams ( McAdams, 1985 ) ; LIWC ( Pennebaker et al., 2015 ) ; Big Five ( McCrae and John, 1992 ) |
| 2. Idiolect | How the user writes at the sentence level, without mechanically copying phrases. Implemented through function-word profile, sentence shape, appraisal style, abstract slot patterns, capitalization, punctuation, formality, and emoji palette. | Martin and White APPRAISAL ( Martin and White, 2005 ) ; Construction Grammar ( Goldberg, 1995 ) |
| 3. Indexical repertoire | Which expressive modes the user can naturally switch into while still sounding like the same person. Implemented through stances, registers, backstage/frontstage range, and speech-genre fluency. | Bakhtin speech-genre theory ( Bakhtin, 1986 ) ; Goffman ( Goffman, 1959 ) |
| 4. Surface modulation | What changes because of app, audience, or context, rather than because the user became a different person. Implemented through length, emoji intensity, self-censoring, disclosure depth, topical focus, and posting rhythm. | Bell audience design ( Bell, 1984 ) |
| Type | What it captures | Theoretical anchor |
|---|---|---|
| Personality trait | Core character attributes | Big Five ( McCrae and John, 1992 ) ; Dark Triad ( Paulhus and Williams, 2002 ) |
| Aspiration | Dreams, goals, aspirational pursuits | Maslow’s hierarchy ( Maslow, 1943 ) |
| Emotional pattern | Recurring emotional dynamics | Uses and Gratifications ( Katz et al., 1973 ) , in its affective branch |
| Identity anchor | Cultural era and tribal belonging, with both overt and covert markers | Social Identity Theory ( Tajfel and Turner, 1979 ) |
| Intimate interest | Body confidence, sensuality, and attraction patterns, anchored to a specific object or aesthetic | Self-presentation ( Goffman, 1959 ) ; Barthes’ punctum ( Barthes, 1981 ) |
| Intellectual curiosity | Hidden learning interests | Self-Determination Theory ( Deci and Ryan, 1985 ) |
| Check | Scoring | What it asks |
|---|---|---|
| Uses relevant preferences | Judge score | Does the output use the user’s relevant preferences when personalization would help? |
| Uses the right amount of personalization | Judge score | Does the output personalize only as much as the situation calls for? |
| Handles social context | Judge score | When other people are involved, does the output handle friends, strangers, groups, and recipients appropriately? |
| Matches user voice | Judge score | When writing for the user, does the output sound like the user and fit the target platform? |
| Avoids disliked topics | Hard check | Does the output avoid topics the user recently disliked, rejected, or asked not to use? |
| Protects private context | Hard check | Does the output avoid revealing sensitive or privacy-flagged information unless clearly needed? |
| Evaluation task | User request | What we check |
|---|---|---|
| Personalized chatbot response | Answer a user query on chatbot using the user’s cross-platform history. | Does the response use the relevant current preference, avoid disliked or contradicted signals, and not sound like a pasted profile? |
| Local recommendation after geo shift | Recommend something local after the user has silently moved cities. | Does the response infer the current city from history, avoid anchoring on the old city, and still fit the user’s general preferences? |
| Personal-fact hallucination probe | Complete a small task that requires a personal fact the user has never shared anywhere in their history. | Does the system notice the missing detail and ask for it, instead of fabricating a plausible value to seem helpful? |
| Understanding hidden persona | Answer a normal user question where a deeper inferred motivation may help. | Does the response serve the user’s deeper motivation without naming the hidden persona or exposing private context? |
| Tracking preference changes | Answer a question after the user’s preference has changed or a short-term intent has expired. | Does the response follow the user’s current stance instead of relying on an older but tempting preference signal? |
| Evaluation task | User request | What we check |
|---|---|---|
| Personalized feed ranking | Rank candidate social media posts for the user’s feed. | Does the system rank the held-out post the user actually engaged with above surface-similar hard negatives and random fillers? |
| @AI directive follow-up | Follow a user’s in-feed instruction to show more or less of a topic. | Does the system still respect the directive after 24 hours, 72 hours, and 7 days, and avoid putting carved-out topics at the top? |
| Hidden-persona recommendation | Recommend feed content that serves a deeper inferred motivation rather than a stated interest. | Does the system surface items aligned with the hidden persona without naming it or exposing private context? |
| Short-term preference lifecycle | Rank recommendations before and after a short-term intent expires. | Does the system use the short-term preference while it is active, then stop relying on it after its expected end time? |
| Evaluation task | User request | What we check | Theoretical anchor |
|---|---|---|---|
| Generic chatbot restraint | Answer a general question where personalization is unnecessary. | Does the system answer naturally without forcing in the user’s preferences irrelevant to the actual question? | Mixed-Initiative interaction ( Horvitz, 1999 ) ; notification interruption science. |
| Sensitive-event restraint | Answer a benign chatbot query shortly after sensitive evidence appears in history. | Does the system avoid mentioning the sensitive event, the planted evidence row, or related private context? | Goffman’s back-stage and front-stage distinction ( Goffman, 1959 ) . |
| Repetitive feed personalization | Recommend repeatedly around the same preference cluster. | After a few allowed repetitions, does the system diversify with new persona-aligned hashtags instead of reusing the same topic? | Skinner’s variable-ratio reinforcement ( Skinner, 1953 ) ; Berlyne’s specific vs. diversive curiosity ( Berlyne, 1960 ) . |
| Repetitive chatbot personalization | Answer varied chatbot questions that repeatedly invite the same preference. | After the preference has already been used several times, does the system stop pivoting every answer back to that same preference? | Skinner’s variable-ratio reinforcement ( Skinner, 1953 ) ; Berlyne’s specific vs. diversive curiosity ( Berlyne, 1960 ) . |
| Do-not-personalize follow-up | Respond in a scenario where personal context would now be awkward or inappropriate. | Does the system avoid leaking forbidden preferences and respect any explicit carve-out from the user? | Goffman’s back-stage and front-stage distinction ( Goffman, 1959 ) ; Horton and Wohl’s parasocial relationship theory ( Horton and Wohl, 1956 ) ; Higgins’s ought-self ( Higgins, 1987 ) . |
| Evaluation task | User request | What we check |
|---|---|---|
| Community voice draft | Draft a social media post aggregated from recent social media activity. | Does the agent draft in the user’s voice from recent social activity, without posting or over-sharing private details? |
| DM inbox digest | Summarize recent social media DMs at chatbot. | Does the agent summarize the right DM threads without sending messages or exposing private details? |
| Cross-app repost adaptation | Repost content from one social media app to another. | Does the agent preserve the source post’s meaning and create exactly one adapted post on the target app? |
| Personalized DM reply | Reply to a friend’s DM on the user’s behalf. | Does the agent send one appropriate reply that addresses the DM and matches the user’s voice? |
| Vague memory refind | Find previously viewed social media content from vague descriptions at chatbot. | Does the agent recover the intended content without creating posts or sending messages? |
| Proactive daily catch-up | Surface a useful daily update without being asked. | Does the agent surface concrete recent updates while avoiding stale, disliked, or overly personal topics? |
| Evaluation task | Backend trigger | What we check |
|---|---|---|
| Close-friend DM update | A close friend sent a DM, and the user has not replied after a meaningful delay. | Does the agent correctly surface the missed friend update without exposing private DM content? |
| Sensitive-event silence | The backend contains recent evidence of a sensitive life event. | Does the agent stay silent instead of proactively bringing up sensitive personal context? |
| Friend-post update | A close friend posted something on social media that matches the user’s interests, and the user has not engaged with it. | Does the agent surface the post only when it is relevant, and do so without sounding invasive? |
| Trending-topic surfacing | A trending social media topic appears, but it may or may not match the user’s interests. | Does the agent surface relevant trends and ignore irrelevant, disliked, or stale ones? |
| Mistake-prevention alert | Cross-platform signals from chatbot, calendar, geolocation, or app history suggest the user may be about to make a mistake. | Does the agent correctly decide whether to warn, ground the warning in concrete evidence, and stay silent when there is no real issue? |
| Idle-moment silence | The user has recent history, but there is no clear reason to interrupt them. | Does the agent stay silent instead of inventing a reason to proactively message the user? |
| Dimension | PersonaMem-v1 | PersonaMem-v2 | PersonaMem-v3 |
|---|---|---|---|
| Data source | 20 fully synthetic users | 1000 fully synthetic users with more comprehensive personas | 200 anonymized real-world users with 4,000,000 engagement histories |
| Explicit vs. implicit | Explicit user preferences | Implicit user preferences | Around 95% implicit user behavior signals |
| Scenarios | Chatbot conversations | Chatbot conversations | Omni-platform , including chatbot, social media recommendation, agentic tasks , and proactiveness |
| Restraint | Personalization | Personalization | Personalization and over-personalization |
| User privacy | No mentioning of user private information | Including personally identifiable information and user-initiated ask-to-forget scenarios | Including psychology-anchored hidden persona and socially inappropriate scenarios |
| Dynamics | Fully synthesized preference updates | Fully synthesized preference updates | Reinforced, emerging, diminishing, bursting, and varied attention shifts from the real world |