cs.AISep 23, 2026

StudentBench: AI and human tutoring yield equivalent GRE learning gains

Authors: Curtis Northcutt, Inaara Hasmani, Kevin Feng, Trevor Khangi, Andreas Plesner, Jonas Mueller

Abstract

Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether large language models (LLMs) produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average. In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations. Together, the two studies clearly separate AI tutors across: (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement. Surprisingly, one AI tutor achieved learning gains equivalent to human tutoring (p = .044) at 918 times lower cost (USD 0.0052 for AI versus USD 4.81 for human, per percentage point gained). For Quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains (all p < .002). To support future research, we open-source the de-identified data collected in our studies.

Figures & tables

Appendix figures & tables33 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. GRADE: Generalizable Reasoning-Aware Dialogue Evaluation for AI Tutors

    May 27, 2026Parth Bhalerao, Jeromy Chang, David Chou +1TutorsGrading

  2. AI-Driven Assessment of Human Tutors: Linking Training Performance to Real-Life Practice

    Jun 17, 2026Danielle R. Thomas, Marie Cynthia Abijuru Kamikazi, Clara Brandt +2TutorsComputerized Adaptive Testing

  3. Knowledge Distillation for Automated AI Tutor Evaluation

    Jul 12, 2026Tahmid Al Hannan, Diego Garcia, Alex Njoroge +2TutorsPedagogical Frameworks