StudentBench: AI and human tutoring yield equivalent GRE learning gains
Abstract
Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether large language models (LLMs) produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average. In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations. Together, the two studies clearly separate AI tutors across: (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement. Surprisingly, one AI tutor achieved learning gains equivalent to human tutoring (p = .044) at 918 times lower cost (USD 0.0052 for AI versus USD 4.81 for human, per percentage point gained). For Quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains (all p < .002). To support future research, we open-source the de-identified data collected in our studies.
Figures & tables
Appendix figures & tables33 assets
Supplementary material from the paper’s appendix.
Appendix
| Section | Condition | Sessions | Pre-test, mean (SD) | P Q | Q P |
| Quantitative | AI | 1,139 | 50.21 (21.16) | 580 | 559 |
| Quantitative | Human | 61 | 45.17 (18.23) | 27 | 34 |
| Quantitative | Control | 91 | 52.99 (21.22) | 45 | 46 |
| Verbal | AI | 1,000 | 46.35 (19.47) | 499 | 501 |
| Verbal | Human | 79 | 41.44 (15.43) | 57 | 22 |
| Verbal | Control | 99 | 47.29 (19.08) | 50 | 49 |
| Primary exclusion reason | Quant | Verbal | Total |
| Eligibility and study-protocol checks | |||
| Pre-test score below eligibility range | 225 | 83 | 308 |
| Pre-test score above eligibility range | 266 | 38 | 304 |
| Insufficient assessment effort | 143 | 124 | 267 |
| Incomplete study | 139 | 34 | 173 |
| Assessment score unavailable | 87 | 51 | 138 |
| AI tutor | Endpoint | Quant | Verbal |
|---|---|---|---|
| Gemini 3.1 Pro (high) | gemini-3.1-pro-preview | 91 | 74 |
| Gemini 3.5 Flash (low) | gemini-3.5-flash | 97 | 80 |
| Gemini 3.6 Flash (low) | gemini-3.6-flash | 89 | 79 |
| Gemini 3.7 Flash (med) | gemini-3.7-flash | 87 | 92 |
| Gemma 4 31B (high) | gemma-4-31b-it | 87 | 72 |
| GPT-5.4 mini (off) | gpt-5.4-mini | 90 | 71 |
| Component | Main-study system templates | Additional guidance in the detailed variants |
| GRE context | Brief content and format overview | Detailed test structure, pacing, scoring conventions and format-specific strategies |
| Diagnosis and sequencing | AI tutor identifies needs and orders concepts | Explicit question-by-question misconception diagnosis, response-time interpretations and prerequisite sequencing |
| Difficulty calibration | AI tutor selects appropriate concepts and practice | Instructions for interpreting easy versus hard mistakes and using supplied expert ratings |
| Practice progression | Generate practice with worked solutions | Vary surface forms and progress from easier to difficult GRE-level practice |
| Conversational teaching | Teach transferable methods; choose explanation and practice | Attempt before hints, guide students to identify mistakes, request explanations and predictions, and connect conceptual understanding to efficient GRE methods |
| Question | Reporting rule |
| Equivalence | TOST at with -SD bounds; is the larger of the two one-sided values. Each AI tutor is tested separately against human tutoring, without correction across tutors; all six passing tutors are highlighted. |
| Differences and associations | Two-sided tests, or the stated omnibus test. For families of related comparisons, we report the Holm-corrected value; the families and individual tests are specified below. |
| Section | Contrast | Estimate | 95% CI | |
| Quantitative | AI control | 6.86 | [4.02,9.69] | |
| Quantitative | Human control | 6.27 | [1.72,10.82] | .007 |
| Verbal | AI control | 5.47 | [2.46,8.47] | |
| Verbal | Human control | 7.52 | [3.17,11.87] | |
| Combined | AI control | 6.15 | [4.08,8.21] | |
| Combined | Human control | 7.04 | [3.88,10.20] |
| Original data | After exclusion | |||
| Result | Estimate | Estimate | ||
| Combined AI human gain | .015 | |||
| Quantitative AI human gain | .028 | .034 | ||
| Student messages per 10 seconds longer reply time | ||||
| Correct practice per 10 more student messages | 0.52 | .002 | 0.47 | .006 |
| Gain per 10 more correct practice problems | 5.62 | 5.77 | ||
| AI tutor | ||||||
| Gemini 3.1 Pro (high) | 165 | 6 | 7 | 7.3% | ||
| Gemini 3.5 Flash (low) | 177 | 2 | 4 | 3.3% | ||
| Gemini 3.6 Flash (low) | 168 | 5 | 4 | 5.1% | ||
| Gemini 3.7 Flash (med) | 179 | 7 | 8 | 7.7% | ||
| Gemma 4 31B (high) | 159 | 3 | 7 | 5.9% | ||
| GPT-5.4 mini (off) | 161 | 1 | 8 | 5.3% |
| Domain | Condition | Gain relative to control [95% CI] |
| Quantitative domains | ||
| Data analysis | AI | 5.7 [1.6, 9.9] |
| Human | 6.8 [0.2, 13.3] | |
| Geometry | AI | 3.1 [ 1.9, 8.2] |
| Human | 1.3 [ 7.3, 9.9] | |
| Arithmetic | AI | 6.2 [2.1, 10.2] |
| Predictor | Scope | Spearman | |
| Cost | Quantitative | 13 | |
| Cost | Verbal | 13 | |
| Cost | Combined | 12 | |
| Reply time | Quantitative | 13 | |
| Reply time | Verbal | 13 | |
| Reply time | Combined | 12 |
| Section | Predictor / outcome | 95% CI | ||
| Quant. | Reply time / student messages | |||
| Quant. | Student messages / correct practice | |||
| Quant. | Correct practice / learning gain | |||
| Verbal | Reply time / student messages | |||
| Verbal | Student messages / correct practice | |||
| Verbal | Correct practice / learning gain |
| Criterion | What the reviewer evaluates |
|---|---|
| A. Relevant concepts | Choosing concepts that address the student’s pre-test errors. |
| B. Concept grouping | Teaching related topics together. |
| C. Concept prioritization | Placing high-impact concepts earlier. |
| D. Time allocation | Dividing time according to errors, difficulty and likely benefit. |
| E. Practice alignment | Matching practice to the student’s errors and skill level. |
| F. Appropriate difficulty | Choosing problems or concepts suitable for what the student knows. |
| Evaluation | Ratings | Pre-tests | Original | Sensitivity |
| Lesson planning | 9,577 | 381 | 9/12 | 8/12 |
| Practice creation/design | 5,265 | 380 | 9/12 | 8/12 |
| All eight shared criteria | 14,842 | 381 | 9/12 | 8/12 |
| Indicator | Detection rule |
|---|---|
| Scaffolding cues | A tutor response following a student message of at least two words contains evaluative language, or shares at least two substantive words with that message and contains a scaffolding cue. |
| Request for explanation | At least one tutor question contains a request for reasoning or explanation. |
| Interactive questions | At least four tutor messages contain a question mark or an invitation to respond. |
| Early attempt request | At least one of the first five tutor messages asks the student to try or attempt a problem. |
| Long solution after a reply | After a student message of at least two words, a tutor message of at least 140 words includes a solution or calculation cue. |
| Reasoning checks | At least two tutor questions contain a reasoning-check or strategy cue. |
| Lesson plan / tutor | Opus 4.8 | Gemini 3.5 Flash | GPT-5.5 | All |
| Minimal / Minimal | 15.23 (9) | 17.78 (10) | 13.56 (62) | 14.27 (81) |
| Minimal / Expanded | 14.81 (27) | 8.50 (27) | 12.65 (36) | 12.06 (90) |
| Intermediate / Intermediate | 11.62 (29) | 2.06 (9) | 11.40 (26) | 10.19 (64) |
| Expanded / Minimal | 8.38 (19) | 2.65 (7) | 23.23 (11) | 11.71 (37) |
| Expanded / Expanded | 8.02 (18) | 11.85 (10) | 8.26 (26) | 8.85 (54) |