AnthroDial: Benchmarking LLM Anthropomorphism in Autonomous Social Interaction
Organizations: Shanghai Institute of Innovation · University of Science and Technology of China · East China Normal University · Anhui University · Shanghai Jiaotong University · Fudan University · University of Melbourne · Zhejiang University · The Hong Kong University of Science and Technology (Guangzhou) · Shanghai Tianyou Software Co., Ltd. · Chabiyue (Shanghai) Information Technology Co., Ltd. · Zhejiang Century Huatong Group Co., Ltd.
Abstract
Large language models (LLMs) are increasingly deployed as social agents, yet credible human-like interaction requires more than fluent responses or persona consistency. Agents must autonomously decide whether, when, and how to communicate while adapting to evolving contexts, goals, and relationships. Existing research, however, lacks a unified approach to enabling, evaluating, and improving such capabilities in continuous, open-ended interaction. We introduce AnthroDial, a unified framework for developing anthropomorphic social agents from three complementary aspects: MindFlow, a lightweight interaction harness that enables autonomous, asynchronous, and adaptive communication through a dynamic Mind Buffer; CAPS-Eval, a theory-grounded framework for evaluating cognitive, affective, and behavioral dimensions of anthropomorphic interaction; and a scalable training paradigm that combines SEEDS for environment expansion with DiAPO for adaptive capability optimization. We further construct evaluation datasets covering everyday communication, game interaction, and long-horizon character interaction. Extensive experiments across diverse models and scenarios demonstrate improved interaction autonomy and naturalness, validate the reliability, discriminativeness, and agreement with human rankings of CAPS-Eval, and confirm the effectiveness of our training paradigm. Together, these components provide a unified framework for developing credible human-like social agents in open-ended interaction.
Figures & tables
| Everyday Chat | Long-Horizon Character | Game Interaction | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Score | Per-Turn | Holistic | Score | Per-Turn | Holistic | Score | Per-Turn | Holistic |
| Claude-4.6 Thinking | 0.9780 | 0.9722 | 0.9844 | 0.9860 | 0.9823 | 0.9888 | 0.8290 | 0.7454 | 0.9135 |
| Gemini-3.5 Flash | 0.9780 | 0.9823 | 0.9747 | 0.9720 | 0.9952 | 0.9494 | 0.8490 | 0.7782 | 0.9200 |
| GPT-5.5 | 0.8820 | 0.8058 | 0.9586 | 0.9820 | 0.9941 | 0.9704 | 0.7010 | 0.5811 | 0.8201 |
| DeepSeek-V4-Flash | 0.9680 | 0.9779 | 0.9589 | 0.9460 | 0.9957 | 0.8966 | 0.7520 | 0.6253 | 0.8782 |
| Qwen3.5-397B-A17B | 0.9410 | 0.9636 | 0.9191 | 0.7600 | 0.8728 | 0.6474 | 0.8360 | 0.7649 | 0.9074 |
| Everyday Chat | Long-Horizon Character | Game Interaction | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Setting | Score | Per-Turn | Holistic | Score | Per-Turn | Holistic | Score | Per-Turn | Holistic |
| Qwen3.5-9B | 0.6410 | 0.7261 | 0.5558 | 0.2120 | 0.3374 | 0.0871 | 0.6620 | 0.5914 | 0.7320 |
| + SFT | 0.9210 | 0.9677 | 0.8737 | 0.7280 | 0.9232 | 0.5338 | 0.8890 | 0.8348 | 0.9429 |
| + SFT + RL (DiAPO) | 0.9801 | 0.9966 | 0.9636 | 0.9930 | 0.9948 | 0.9904 | 0.9926 | 0.9981 | 0.9871 |
| + SFT + RL (GRPO) | 0.9731 | 0.9962 | 0.9501 | 0.9820 | 0.9978 | 0.9664 | 0.9879 | 0.9933 | 0.9825 |
| Everyday Chat | Long-Horizon Character | Game Interaction | |||||||
| Average score | Pass rate | |||||
|---|---|---|---|---|---|---|
| Variant | Score | Per-Turn | Holistic | ACC@85 | ACC@90 | ACC@95 |
| GRPO | 0.9731 | 0.9962 | 0.9501 | 0.9800 | 0.8800 | 0.7200 |
| + CDT only | 0.9758 | 0.9964 | 0.9551 | 1.0000 | 0.9600 | 0.6800 |
| + ZPD only | 0.9751 | 0.9974 | 0.9529 | 1.0000 | 0.9200 | 0.7000 |
| + CDT + ZPD | 0.9801 | 0.9966 | 0.9636 | 1.0000 | 0.9800 | 0.7600 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Subset | Evaluation size | Category distribution (scenario counts) |
|---|---|---|
| Everyday Chat | 12 categories 50 scenarios / bindings 46 persona cards 100 role cases | Work communication: 5; Daily chat: 5; Emotional support: 5; Interests: 5; Relationship development: 4; Pets and family life: 4; Family and parenting: 4; Sports: 4; Finance: 4; News: 4; Gaming: 3; Technology: 3. |
| Game Interaction | 4 categories 59 scenarios / bindings 95 persona cards 118 role cases | Gameplay and mechanics: 32; Accounts, assets, and systems: 15; Gaming: 7; Social interaction and daily life: 5. |
| Long-Horizon Character | 5 categories 50 scenarios / bindings 100 persona cards 100 role cases | Daily chat: 22; Interests: 13; Technology and finance: 6; Daily life and emotions: 5; Sports and competition: 4. |
| Subset | Hard constraints | Per-turn | Holistic |
|---|---|---|---|
| Everyday Chat | 10 | 5 / 20 | 5 / 22 |
| Game Interaction | 14 | 11 / 29 | 9 / 20 |
| Long-Horizon Character | 5 | 5 / 11 | 5 / 18 |
| ID | Constraint | Sev. | Operational evidence |
|---|---|---|---|
| L0-01 | AI identity exposure | F | Mentions being an AI, language model, or unable to act like a real person. |
| L0-02 | Text-message medium | F | Claims to send or receive images, voice, video, files, screenshots, or links in a text-only chat. |
| L0-03 | Non-co-presence | F | Claims to see, sit beside, or physically point to the partner when the scenario is remote chat. |
| L0-04 | Output contract | F | Missing or malformed role, response, content, or timestamp fields. |
| L0-05 | Hard repetition | F | Exact copy of an earlier message or the partner’s immediately preceding message. |
| L0-06 | Persona hard facts | F | Contradicts explicit age, gender, occupation, relationship, location, or relationship status. |
| ID | Dimension | Wt. | Checkbox criteria |
|---|---|---|---|
| Per-turn | |||
| D1 | Chat naturalness | 13 | (1) Colloquial private-message style. (2) No formal connectors. (3) No service or assistant tone. (4) Chinese WeChat habit. |
| D2 | Information density | 9 | (1) Length fits the context. (2) No over-explaining. (3) Each message has a useful intention. (4) No empty filler bursts. |
| D3 | Persona expression | 14 | (1) Age-language match. (2) Stable personality. (3) Natural catchphrases and punctuation. (4) Occupation/education fit. (5) Role-specific social style. |
| D4 | Knowledge and action boundaries | 8 | (1) No excess expertise. (2) Natural uncertainty on weak topics. (3) No unauthorized actions. (4) Responds from life experience rather than fake authority. |
| D5 | Rhythm and segmentation | 8 | (1) Appropriate segmentation count. (2) Each segment is readable. (3) Segment order and reply rhythm are plausible. |
| ID | Constraint | Sev. | Operational evidence |
|---|---|---|---|
| L0-01 | AI identity exposure | F | Discloses being an AI or language model. |
| L0-02 | Text-only medium | F | Claims to send or receive non-text content such as images, voice, video, or files. |
| L0-03 | Non-co-presence | F | Suddenly places both participants together unless the scenario explicitly permits it. |
| L0-04 | Output contract | F | Missing or malformed JSON, role, response, content, or timestamp fields. |
| L0-05 | Hard repetition | F | Exactly copies a historical message. |
| L0-06 | Persona facts | F | Contradicts explicit gender, age, occupation, relationship, location, or relationship status. |
| ID | Dimension | Wt. | Checkbox criteria |
|---|---|---|---|
| Per-turn | |||
| D1 | Game-chat naturalness | 6.5 | (1) Use casual, concise language. (2) Use particles naturally when present; terse replies without particles also pass. (3) Avoid formal written connectors. |
| D2 | Information density | 4.5 | (1) Match length to the relationship, situation, and emotion. (2) Avoid over-explaining. |
| D3 | Persona expression | 7 | (1) Match language to the persona’s age. (2) Maintain personality traits. (3) Use catchphrases naturally when present; their absence also passes. |
| D4 | Rhythm and segmentation | 4 | (1) Use an appropriate number of message segments. (2) Keep each message independently readable. |
| D5 | Knowledge boundaries | 4 | (1) Avoid expertise beyond the persona. (2) Express uncertainty when a knowledge gap arises; no such gap also passes. |
| ID | Constraint | Sev. | Operational evidence |
|---|---|---|---|
| L0-01 | AI identity exposure | F | Discloses being an AI or language model. |
| L0-02 | Safety/ethics | F | Gives clearly illegal, dangerous, or harmful advice. |
| L0-03 | Hard repetition | F | Exactly copies a historical message. |
| L0-04 | Output contract | F | Missing or malformed JSON, role, response, content, or timestamp fields. |
| L0-05 | Persona facts | F | Contradicts explicit gender, age, occupation, relationship, or location. |
| ID | Dimension | Wt. | Checkbox criteria |
|---|---|---|---|
| Per-turn | |||
| D1 | Interpretable expression | 10 | (1) Make intent, attitude, or reference recoverable from the message and local context; fragments, typos, interruptions, and short acknowledgments may pass. (2) Keep text usable as dialogue; formality, length, and incomplete sentences alone are not failures. |
| D2 | Autonomous responses | 10 | (1) Fail only after three consecutive replies neither engage the partner nor provide natural listening signals; brief acknowledgments during narration pass. (2) Fail only after three consecutive substantive replies echo or paraphrase the partner without adding information or stance; ordinary restatement and listening signals pass. |
| D3 | Appropriate interaction wording | 10 | (1) Fail only when three consecutive replies keep pushing questions, proposals, or plans after an answer, disengagement, or topic departure. (2) Fail when three consecutive visible replies repeat the fixed agree/restate–elaborate–question/plan structure. (3) Respect an explicit refusal, stop request, or topic boundary immediately; no explicit boundary passes. |
| D4 | Grounded emotional wording | 10 | (1) Do not interpret emotion as the opposite of the latest explicit cue; emotional expression is optional. (2) Fail only when three consecutive visible replies exactly mirror emotional intensity without independent response or change; natural resonance passes. |
| D5 | Local template avoidance | 10 | (1) Fail when three substantive replies preserve a near-identical functional sequence or syntactic frame across different inputs; isolated similarity and listening signals pass. (2) Fail when three consecutive substantive replies strongly overlap without new information or stance; one repetition, two retransmissions, and listening signals may pass. |
| Stage | Operation |
|---|---|
| Case construction | Load a binding and create two cases, one with the tested model as and one with it as . |
| Dialogue generation | Instantiate both personas with the same scenario, opener, virtual start time, duration horizon, and single-draft scheduler; save dialogue and raw model-call traces. |
| L0 gate | Check deterministic identity, medium, co-presence, structure, and repetition failures; mark fatal cases invalid before soft scoring. |
| Per-turn scoring | For each evaluated-role turn, call the judge with the target turn and up to preceding turns; compute D1–D5 checkbox scores. |
| Holistic scoring | Call the judge once with the full dialogue and evaluated role marker; compute E1–E5 checkbox scores. |
| Aggregation | Aggregate Score and paired ACC@ ; retain validity rates, case details, dimension and category breakdowns, and all-checkbox-pass diagnostics. |
| Model | Human Annotators | Qwen3.5-397B-A17B | DeepSeek-V4-Flash |
|---|---|---|---|
| Claude-4.6 Thinking | 1.0000 (1) | 0.9780 (1) | 0.9830 (1) |
| Gemini-3.5 Flash | 0.9560 (2) | 0.9780 (1) | 0.9790 (2) |
| Qwen3.5-397B-A17B | 0.9410 (3) | 0.9410 (4) | 0.9650 (4) |
| DeepSeek-V4-Flash | 0.9340 (4) | 0.9680 (3) | 0.9790 (2) |
| GPT-5.5 | 0.8900 (5) | 0.8820 (5) | 0.9140 (5) |
| Qwen3.5-9B | 0.6260 (6) | 0.6410 (6) | 0.7390 (6) |
| Subset | Scenario cards | Persona cards | SFT samples |
|---|---|---|---|
| Everyday Chat | 184 | 290 | 7,774 |
| Game Interaction | 18 | 36 | 430 |
| Long-Horizon Character | 178 | 314 | 22,645 |
| Category | Scenario cards | Scenario share | SFT samples | Training share |
|---|---|---|---|---|
| Interests | 38 | 20.65% | 1,662 | 21.38% |
| Daily chat | 15 | 8.15% | 440 | 5.66% |
| Emotional support | 21 | 11.41% | 1,072 | 13.79% |
| Technology | 15 | 8.15% | 618 | 7.95% |
| Family and parenting | 13 | 7.07% | 491 | 6.32% |
| Relationship development | 3 | 1.63% | 116 | 1.49% |
| Everyday Chat | Long-Horizon Character | Game Interaction | |||||||
| Setting | Score | Per-Turn | Holistic | Score | Per-Turn | Holistic | Score | Per-Turn | Holistic |
| Baseline dialogue methods | |||||||||
| ECP | 0.7140 | 0.8104 | 0.6178 | 0.0880 | 0.1335 | 0.0431 | 0.7310 | 0.6798 | 0.7822 |
| HumanLM | 0.6070 | 0.6405 | 0.5736 | 0.1880 | 0.2684 | 0.1079 | 0.7700 | 0.7701 | 0.7692 |
| OSIM | 0.3040 | 0.2035 | 0.4043 | 0.0720 | 0.1025 | 0.0407 | 0.3570 | 0.4665 | 0.2482 |
| SEEDS–DiAPO pipeline (ours) | |||||||||