cs.CLOct 7, 2026

Alice: A Large-Scale German Benchmark for Rubric-Based Multi-Dimensional Automatic Short Answer Scoring

Authors: Zhifan Sun, Sebastian Gombert, Jannik Lossjew, Tobias Wyrwich, Berrit Katharina Czinczel, David Bednorz, Marcus Kubsch, Knut Neumann, +1 more

Organizations: DIPF | Leibniz Institute for Research and Information in Education · IPN | Leibniz Institute for Science and Mathematics Education · Umeå University · Computer Science Department & Studiumdigitale, Goethe University Frankfurt

Abstract

Automatic Short Answer Scoring (ASAS) is central to NLP for Education. However, openly available benchmarks remain scarce, and existing datasets largely address how well students answer a question directly rather than how well they master underlying concepts (knowledge elements) such as thermal energy or epistemic activities (skills) such as reasoning or claim. To address this gap, we introduce Alice, a large-scale, rubric-based German ASAS dataset that is pedagogically aligned and comprises three subtasks: (i) learning performance (Alice-LP), (ii) knowledge elements (Alice-KE), and (iii) skills (Alice-SK). We further formulate rubric-based ASAS as a rubric-retrieval task and benchmark the dataset with a range of language models, from encoder-only models to lightweight LLMs. We also benchmark the dataset with zero-shot prompting via LLMs and a standard classification baseline. The experiments show that LLMs, in particular, struggle to score knowledge elements and skills in the zero-shot setting. They also indicate that rubric text is often useful, especially for Alice-KE and Alice-SK, while on Alice-LP gains over sample-solution-focused inputs are more modest and vary by model and input format.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Rubric Spans are Label Representations: Joint LLM Encoding for Short Answer Scoring

    Oct 7, 2026Zhifan Sun, Sebastian Gombert, Fabian Zehner +3Rubric-Based Scoring

  2. Elmes*: Automated Construction of Fine-Grained Evaluation Rubrics for Large Language Models in Long-Tail Educational Scenarios

    Jun 4, 2026Tao Liu, Ye Lu, Ruohua Zhang +4Large Language Model EvaluationPedagogical Frameworks

  3. When Rubrics Change: Cross-Rubric Generalization for Critical Thinking Essay Scoring

    Jul 15, 2026Nischal Ashok Kumar, Payu Wittawatolarn, Sana Kang +7Automated Essay ScoringTask-Specific Rubrics