cs.CLOct 6, 2026

Wiki-Talkie: Multilingual Benchmarking of Persona-Based Agents on Real-World Discussions

Authors: Dennis Fucci, Andrea Bacciu, Dong Liu, Weronika Łajewska, Saab Mansour

Organizations: Amazon

Abstract

LLMs are increasingly deployed as autonomous agents in social environments, making it critical to study their ability to faithfully simulate human interactions. Central to this is grounding agents in realistic user personas, yet existing datasets rely on fictional personas and are limited to a handful of languages, lacking the empirical grounding necessary to evaluate behavioral fidelity across diverse populations. We introduce Wiki-Talkie, a multilingual dataset of real-world conversations from Wikipedia Talk pages across five languages spanning two language families: Germanic (German, English) and Romance (Spanish, French, Italian), paired with personas derived from real user communities and encompassing sociodemographic attributes, self-descriptions, and behaviorally grounded interaction traits. Using Wiki-Talkie, we evaluate agent interactional behavior on a next-turn generation task across various persona conditioning strategies. Our evaluation assesses whether agents collectively reproduce the distributional behavioral patterns observed in human discussions. Results show that user's comment history exemplifying interaction behavior consistently outperforms explicit persona information. In addition, models systematically underproduce negative or extreme sentiments, while over producing references and suggestions, revealing biases toward agreeableness and positivity. Crucially, these patterns hold robustly across languages, with small cross-lingual differences.

Figures & tables

Appendix figures & tables31 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents

    May 13, 2026Harshita Chopra, Kshitish Ghate, Aylin Caliskan +3User SimulationLarge Language Model Agents

  2. WorldBench: Culturally Grounded Benchmark for Multilingual Agents

    Sep 1, 2026Leonardo Ranaldi, Sherrie Shen, Jushi Kai +1Multilingual AgentsAgentic Benchmarks

  3. PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models

    May 16, 2026Wenlong Shi, Jianxun Lian, Mingqi Wu +5Role-Playing AgentsPersona Consistency