cs.CLOct 3, 2026

AI-Enabled Quality Assurance for Multiple-Choice Assessment Items

Authors: Steven Moore, Nicholas Diana

Organizations: George Mason University · Colgate University

Abstract

Generating multiple-choice questions is increasingly scalable, but establishing their assessment quality remains difficult. We present a focused narrative review of automated item-writing flaw detection, revision, psychometric screening, and NLP benchmark auditing. Database searches, citation retrieval, and nominated sources yield fourteen research reports reviewed in full text. We distinguish surface checks from content-sensitive judgments and map a 19-criterion rubric to detection methods and reported evidence. High label-level accuracy often coexists with weak positive case detection, while rubric definitions and reference standards vary. Revision evidence is mixed, and the associations reported in prior work do not establish the effects of repair. We propose evaluating quality assurance as a sequence of independently validated decisions, with criterion-specific reporting, calibrated human review, and outcome-based assessment of revisions.

Explore similar work

CardsList
  1. Auditing MCQA Benchmarks through Probability Landscapes

    Aug 31, 2026Minsoo Song, Chanjun ParkBenchmark AuditingLLM Auditing

  2. Automated item evaluation: Predicting item acceptance and rejection using LLM-generated critiques

    Aug 6, 2026Hotaka Maeda, Yikai LuEducational AssessmentText Classification

  3. The Architect, the Adversary, and the Judge: Closed-Loop Generation of Standards-Aligned Assessment Items at Scale

    Aug 26, 2026Wenhui Chen, Ziyao Lin, Jianlin Chen +2Educational AssessmentLLM-as-a-Judge