Agreement Is Not Validity: Cross-Model LLM Consensus in Diagnosing Student Failure Modes in K-12 Math Tutoring Dialogue
Organizations: College of Computing and Information Science, Cornell University, Ithaca, USA. · College of Connected Computing, Vanderbilt University, Nashville, USA. · Third Space Learning, Swindon, UK.
Abstract
In K-12 mathematics tutoring, student-tutor dialogue provides rich evidence of learners' problem-solving processes and sources of difficulty. Learning analytics research increasingly relies on large language models (LLMs) to extract such information from dialogue for a variety of downstream tasks, including knowledge tracing, behavioral modeling, and diagnosis of student reasoning errors. However, the validity of these model-generated interpretations remains insufficiently understood. In this exploratory study, we examine the validity of LLM classifications of five student failure modes in mathematics tutoring dialogue using an operational diagnostic codebook: uncertainty, misattribution, operator selection, conceptual gap, and procedural slip. Across models, human-LLM agreement was moderate (kappa = .524-.597), while cross-model agreement was substantially higher (kappa = .755-.781; alpha = .769). These findings show that cross-model agreement can create a misleading appearance of correctness, challenging the assumption that consensus among LLMs constitutes evidence of valid learner interpretation. For learning analytics, the implication is clear: scalable labeling is useful only if the inferred constructs are valid, and model consensus cannot substitute for independent evidence of that validity.
Figures & tables
| Code | Failure Mode | Definition | Example |
|---|---|---|---|
| UNC | Uncertainty | Student expresses uncertainty or reports a technical issue | “I don’t know”; “Um…” |
| MIS | Misattribution | Student uses numbers or context from a different problem | Answers another question than asked |
| OP | Operator Selection | Student chooses the wrong type of operation for this problem | Multiplies when division is needed |
| CON | Conceptual Gap | Reasoning reveals a wrong underlying mathematical idea | Believes |
| PROC | Procedural Slip | Correct strategy, wrong execution | Correct method with arithmetic error |
| NA | Not Applicable | Exclusion flag (not a failure mode)—item cannot be diagnosed | Inaudible response; lack of information |