cs.AIOct 4, 2026

The GenAI4IDN Benchmark 3.0 - a Public Tool to Assess Generative AI Tools for the Design of Interactive Digital Narratives

Authors: Hartmut Koenitz, Jonathan Barbara, Mirjam Palosaari Eladhari

Organizations: Södertörn University, Alfred Nobels allé 7, 141 89 Huddinge, Sweden · Saint Martin’s Institute of Higher Education, Hamrun, Malta · Stockholm University, SE-106 91 Stockholm, Sweden

Abstract

This paper presents GENAI4IDN Benchmark 3.0, the third iteration of an evaluation framework to assess Generative AI tools for creating Interactive Digital Narratives (IDNs). Moving beyond manual testing, this iteration introduces AI-assisted evaluation through a publicly accessible web application (https://genai4idn.com), enabling the community to run benchmarks on demand, add new models, and propose new tasks. The revised evaluation framework is "blinded" to avoid model-bias, and can handle complex media such as music, videos, and full IDNs that previously required human raters. A significant addition - responding to concerns raised during ICIDS 2025 - is the addition of fact-checking and bias detection with dedicated tasks and rubrics, validated by human raters with lived experience in the depicted contexts. Findings from a diverse range of models report on maturing creative capabilities while observing runaway thinking and overzealous safety filters as limitations. Fact-checking reliably caught subtle historical inaccuracies, anachronisms, and fabricated claims while the bias rater consistently exposed structural assumptions, tropes, and marginalized group erasures.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. 'Your AI Text is not Mine': Redefining and Evaluating AI-generated Text Detection under Realistic Assumptions

    Jun 3, 2026Nils Dycke, Marina Sakharova, Nico Daheim +1Machine-Generated Text DetectionSri Lanka

  2. AI-Assisted Systematization for Evaluating GenAI Systems

    May 25, 2026Dhruv Agarwal, Emily Sheng, Chad Atalla +6Generative Artificial IntelligenceArtificial Intelligence Evaluation