cs.CYJun 1, 2026

Question Type, Cognitive Load, and CEFR Alignment: Evaluating LLM-Generated EFL Grammar Drill Exercises

Authors: Steve WoollastonBrendan FlanaganYuko ToyokawaHiroaki Ogata

Organizations: Kyoto University, Japan · Ritsumeikan University, Japan

Abstract

This study evaluates the pedagogical viability of LLM-generated English as a Foreign Language (EFL) learning content. Utilising log data from Japanese junior high school students practicing on a grammar drilling application, we analysed how different question modalities impact student performance and whether theoretical localised CEFR difficulty tiers accurately predict empirical task difficulty. Results reveal a clear performance hierarchy: multiple-choice questions carried the lowest cognitive load, cloze tasks posed the greatest barrier to active recall, and drag-and-drop exercises incurred the heaviest time penalties. Furthermore, learner data validated the CEFR-J grammar framework, showing a steady decline in accuracy and increased response times as proficiency levels advanced. These findings demonstrate that LLMs can successfully generate learning content, while highlighting the need for developers to strategically sequence question modalities to transition learners from passive recognition to active linguistic production.

Explore similar work

CardsList
  1. The Crutch or the Ceiling? How Different Generations of LLMs Shape EFL Student Writings

    Apr 16, 2026Hengky Susanto, David James Woo, Chingyi Yeung +2Tutors