cs.CLAug 24, 2026

What Reaches Expert Review? Representation, Structural Screening, and Candidate-Form Dependence in AI-Assisted Item Development

Authors: Christopher Brooks

Organizations: School of Information, University of Michigan

Abstract

Between AI-assisted item generation and expert review sits a computational evaluator whose decisions are usually treated as technical preliminaries. Yet representation, structural reduction, and selection policy determine which items and evidence psychometricians ever receive. Across two linked in-silico studies of 32,000 selected Big Five items, we followed fixed source populations from semantic representation through structural evaluation and candidate-form construction. Broad agreement in semantic geometry concealed consequential local differences: identical wording acquired different construct evidence, different items survived, and intended attributes could disappear even as community correspondence improved. These sensitivities also differed across generated source populations. At the final review boundary, both eligibility policies filled every content cell in every evaluable form, yet they presented different wording. Across embedding configurations, inclusive primary forms shared a median of only 6 of 40 items, reflecting the total downstream consequence of changing representation across structural evidence and ranking. The apparent stability of global summaries and complete forms therefore concealed instability in the content reaching psychometricians. The computational evaluator is not neutral infrastructure between generation and expertise; it is an inspectable and revisable part of measurement design.

Explore similar work

CardsList
  1. The Architect, the Adversary, and the Judge: Closed-Loop Generation of Standards-Aligned Assessment Items at Scale

    Aug 26, 2026Wenhui Chen, Ziyao Lin, Jianlin Chen +2Educational AssessmentLLM-as-a-Judge

  2. AI-Enabled Quality Assurance for Multiple-Choice Assessment Items

    Oct 3, 2026Steven Moore, Nicholas DianaAutomated EvaluationEducational Assessment

  3. Automated item evaluation: Predicting item acceptance and rejection using LLM-generated critiques

    Aug 6, 2026Hotaka Maeda, Yikai LuEducational AssessmentText Classification