cs.CLMay 8, 2026

NARRA-Gym for Evaluating Interactive Narrative Agents

Authors: Yue HuangYuchen MaJiayi YeWenjie WangZipeng LingXingjian HuYuexing HaoZichen Chen+9 more

Organizations: University of Notre Dame · 2LMU Munich · 11Munich Center for Machine Learning · 3Independent Researcher · University of Pennsylvania · 5Lehigh University · 6Massachusetts Institute of Technology · 7Bake AI · 8UC Santa2026 Barbara · 9Stanford University · University of Washington

Abstract

Interactive narrative tasks require LLMs to sustain a coherent, evolving story while adapting to a user over multiple turns. However, suitable benchmarks for this setting are limited: existing evaluations often focus on static prompts, isolated story generations, or post-hoc ratings, and therefore miss whether models can jointly manage story generation, long-context state and pacing, character simulation, empathic personalization, and story-grounded artifacts. We introduce NARRA-Gym, an executable evaluation environment that turns a sparse emotional seed into a complete interactive story episode and logs the full model-in-the-loop trajectory, including story construction, memory updates, planning, pacing interventions, and optional artifact synthesis. We evaluate nine frontier LLMs using a controlled LLM-as-judge sweep over eight benchmark personas and a human evaluation in which participants rate customized model outputs. Our results show substantial variation across models, personas, and evaluation dimensions: models that produce fluent stories can still fail on robustness, user experience, or resistance-sensitive personalization. These findings suggest that interactive narrative offers a useful benchmark for evaluating long-horizon, user-adaptive LLM behavior beyond isolated story quality.

Explore similar work

CardsList
  1. The Story Shapes the Agent: Narrative Priors in LLM Behavior

    Jul 20, 2026Yixuan Wang, James Lester, Shashank SrivastavaPersonaNarratives