Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors reward the student's success on the tutored problem with the tutor's words still in context. The reward is then easiest to raise by telling the student the answer, and a tuned penalty is needed to reduce telling. Drawing on learning sciences, we introduce a masked near-transfer post-test: the student is tested on an unseen variant of the tutored problem with the tutor's utterances masked, so the reward can rise only through what the student wrote in its own turns. This discourages cognitive offloading by the student and allows the continuous penalty to be replaced by two binary reward gates (factual correctness of tutor response, no solution handover). A leave-one-out ablation shows that the learning-gain reward on its own does not separate teaching from telling: the gates reduce solution handover while the near-transfer post-test improves out-of-domain transfer. Using these reward designs we develop Eduardo, a multi-turn RL recipe for training LLM tutors, and use it to train 4B, 9B, 14B and 27B models from two distinct LLM architectures. Our post-trained Eduardo-27B model matches Gemini-3.1-Pro on MathTutorBench and Claude Opus 4.8 on TutorMoments at 2.4-6.2x fewer thinking tokens than frontier models, which matters for interactive tutoring. Without being named in the reward, the model more than doubles its use of the push-for-justification teacher move while support fading (e.g., assigning independent work), whose payoff lies beyond a single-problem dialog episode, is trained out. We open-source our training environment, an 8,671-problem near-transfer dataset, and trained models for further development.
Figures & tables
Figure 1: Mistake diagnostic versus pedagogy on MathTutorBench . Our models (Eduardo) improve pedagogy over their base at every size, while preserving diagnostic accuracy.
Training environment ( n=494 , full dialog)
MathDial ( n=500 , middle turn)
Eedi ( n=500 , middle turn)
Model
Δtransfer↑
Δsame
Leak ↓
Ped. ↑
Fact. ↑
RM ↑
Win ↑
Ped. ↑
Fact. ↑
Think ↓
RM ↑
Win ↑
Ped. ↑
Fact. ↑
Think ↓
Base models
Qwen3.5-4B ∗†
+0.10
+0.11
17%
2.26
2.88
2.08
30.2%
2.65
2.90
938
3.56
38.7%
2.91
3.27
1028
Qwen3.5-9B ∗†
+0.22
+0.35
23%
3.29
4.20
3.02
38.0%
2.94
3.17
1172
6.04
56.7%
3.04
3.70
1404
Qwen3-14B †
+0.19
+0.20
98%
1.24
3.45
7.01
70.2%
3.05
3.84
461
7.91
66.3%
2.96
4.34
357
Qwen3.8-27B
+0.22
+0.44
9%
4.20
4.77
8.34
82.4%
3.67
4.12
501
9.99
84.3%
3.27
4.50
258
Table 1: Main results. Training environment performance and zero-shot out-of-domain evaluation on MathDial and Eedi. Ped. and Fact. is the 1–5 Gemini-3.1-Pro whole-dialog score; Think is mean thinking tokens per turn. ∗ / † : >15% empty turns / truncated dialogs; environment metrics are unreliable (Table 8 ). ‡ Self-judged.
Environment ( n=494 )
MathDial ( n=500 , middle turn)
Eedi ( n=500 , middle turn)
Configuration
Δtransfer↑
Δsame
Leak ↓
Ped. ↑
Fact. ↑
Think ↓
RM ↑
Win ↑
Ped. ↑
Fact. ↑
Think ↓
RM ↑
Win ↑
Ped. ↑
Fact. ↑
Think ↓
Base Qwen3.5-4B (no RL) ∗†
+0.10
+0.11
17%
2.26
2.88
596
2.08
30.2%
2.65
2.90
938
3.56
38.7%
2.91
3.27
1028
Eduardo-4B (full recipe)
+0.25
+0.37
31%
3.66
4.19
544
7.92
78.0%
3.44
3.79
456
8.83
72.4%
3.11
4.15
184
w/o Binary reward gates
+0.23
+0.38
61%
2.80
3.43
398
6.10
65.0%
2.80
3.22
345
8.18
74.0%
3.06
3.94
152
w/o Masking ∗†
+0.24
+0.39
13%
3.18
3.99
718
6.48
65.0%
3.15
4.00
1597
7.48
59.0%
3.13
4.34
1041
w/o Near-Transfer
+0.23
+0.40
47%
2.72
3.30
761
5.25
56.0%
2.99
3.39
776
7.36
65.0%
3.03
4.07
512
Table 2: Leave-one-out ablation (4B). While no ablated condition changes Δtransfer measurably, the conditions determine how the gain is achieved. ∗ / † : >15% empty turns / truncated dialogs; environment metrics are unreliable (Table 8 ).
Plain prompt ( n=100 )
Evaluation-aware prompt ( n=100 )
Tutor
Scaff. ↑
Rigour ↑
Avoids OH ↑
Think/turn ↓
Ratio ↓
Tok./turn ↓
Scaff. ↑
Rigour ↑
Avoids OH ↑
Think/turn ↓
Ratio ↓
Tok./turn ↓
Qwen3.8-27B
0.60
0.31
0.52
∼ 970
6.2 ×
∼ 43
0.90
0.60
0.81
∼ 1099
8.2 ×
∼ 46
Gemini 3.5 Flash
0.94
0.31
0.75
886
5.7 ×
60
0.92
0.83
0.94
886
6.6 ×
54
Gemini 3.6 Flash
0.85
0.40
0.71
602
3.9 ×
44
0.94
0.79
0.91
843
6.3 ×
41
Claude Opus 4.8
0.83
0.48
0.67
∼ 369
2.4 ×
∼ 84
0.98
0.92
0.97
∼ 774
5.8 ×
∼ 88
Eduardo-27B (ours)
0.96
0.44
0.76
∼ 155
1.0 ×
∼ 28
0.98
0.67
0.86
∼ 135
1.0 ×
∼ 27
Table 3: Out-of-domain generalization on TutorMoments dialogs (with independent Gemini-3.7-Flash student and Gemini-3.1-Pro judge) . Judge scores (Scaffolding, rigour, Avoids OH) are in [0,1] . Think/turn is thinking tokens per turn, Ratio is the overhead relative to Eduardo-27B, and Tok./turn is visible output tokens per turn. ( ∼ ) denotes estimates from visible reasoning traces (3.28 chars/token). Under a plain prompt, Eduardo-27B achieves top Scaffolding and Avoids OH scores with 2.4–6.2 × fewer thinking tokens than baselines.
Figure 2: Two routes to the prompted move distribution. (a) JS distance of each model’s move distribution from human tutors ( x ) and from the centroid of the evaluation-aware frontier models ( y , leave-one-out) (b) Per-move shares under the plain prompt, base → Eduardo , against human tutors (diamonds) and the prompted range. n=100 moments per LM cell.
Figure 3: Model size versus MathTutorBench tutoring quality . Represents the mean over the three student-understanding and four pedagogical axes. One recipe, one hyperparameter set: our models improve over their base at every size, and the 27B model approaches top-tier models (0.79 vs. 0.80 for Gemini 3.1 Pro) with fewer thinking tokens.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Value
Optimization
Algorithm
DPPO ( Qi et al., 2026 )
Group size G
16
Mini-batch size
8 (2 asynchronous off-policy steps)
Optimizer
AdamW ( β1=0.9 , β2=0.95 , ϵ=10−15 )
Learning rate
1×10−6 (constant)
Appendix
Table 4: Training hyperparameters. Identical across all model sizes.
Train ( n=8,238 )
Test ( n=433 )
Source
Transfer
Corr.
Source
Transfer
Corr.
Mean
0.289
0.280
0.648
0.283
0.280
0.675
Median
0.250
0.188
0.250
0.250
Std
0.246
0.246
0.240
0.237
Range
[0, 1]
[0, 1]
[0, 1]
[0, 1]
Appendix
Table 5: Dataset solve-rate statistics . Solve rates measured with k=16 samples from frozen Llama-3.1-8B-Instruct.
Source problem
Near-transfer variant
What is the sum of the prime factors of 85,085? Answer: 53 ppre : 0.250
What is the sum of the prime factors of 34,034? Answer: 50 ppre : 0.188
What is the sum of all integer solutions to ∣n∣<∣n−3∣<9 ? Answer: −14ppre : 0.312
What is the sum of all integer solutions to ∣m∣<∣m−5∣<12 ? Answer: −18ppre : 0.188
Given that m is a root of the equation x2−2x−7=0 , find m2−2m+1 . Answer: 8 ppre : 0.375
If r is a root of the equation x2−5x−4=0 , determine the value of r2−5r+10 . Answer: 14 ppre : 0.375
A triangle has an even perimeter, and two of its sides are 5 and 2,008. How many such triangles exist? Answer: 4 ppre : 0.250
Two sides of a triangle have lengths 7 and 2,024. If the perimeter of the triangle is an even integer, how many such triangles are possible? Answer: 6 ppre : 0.188
Appendix
Table 6: Near-transfer pair examples from the test set. Each transfer variant preserves the solution method while changing numbers or surface context, yielding a different answer.
Δtransfer
Δsame
Leak rate
Model
Mean
95% CI
Mean
95% CI
Mean
95% CI
Base models
Qwen3.5-4B ∗†
+0.10
[0.06,
0.13]
+0.11
[0.08,
0.15]
16.8%
[13.6,
20.2]
Qwen3.5-9B ∗†
+0.22
[0.18,
0.26]
+0.35
[0.31,
0.39]
22.7%
[19.0,
26.3]
Qwen3-14B †
+0.19
[0.15,
0.22]
+0.20
[0.17,
0.24]
97.8%
[96.4,
99.0]
Qwen3.8-27B
+0.22
[0.18,
0.25]
+0.44
[0.40,
0.48]
8.9%
[6.5,
11.5]
Appendix
Table 7: Bootstrap 95% CIs for Δtransfer , Δsame , and leak rate ( n=494 problems, 10,000 resamples). ∗ / † as in Table 1 .
Model
Avg turns
Empty%
Think tok/turn
Seq. trunc.%
EoC rate
Qwen3.5-4B (base) ∗†
4.40
70.0%
596
31.2%
68.8%
Qwen3.5-9B (base) ∗†
4.74
19.5%
1504
15.2%
84.8%
Qwen3-14B (base) †
6.35
8.1%
1102
26.1%
59.5%
Qwen3.8-27B (base)
7.05
1.3%
427
3.8%
82.0%
Eduardo-4B
8.18
3.7%
544
10.7%
69.6%
Eduardo-9B
7.87
2.1%
496
6.7%
76.3%
Appendix
Table 8: Truncation diagnostics per model on the 494-problem test set. Avg turns : mean tutor turns per dialog. Empty% : fraction of turns where reasoning exceeded the per-turn cap and produced no visible output. Think tok/turn : mean thinking tokens per tutor turn. Seq. trunc.% : fraction of conversations cut short by the 16,384-token sequence budget. EoC rate : fraction of dialogs ending with the tutor’s end-of-conversation token (vs. hitting turn or token limits). Values >15% in bold . Rule used in all tables: ∗ Empty% >15% , † Seq. trunc.% >15% ; for flagged rows Δ and Leak are understated and judge scores are depressed.
Score
Description
5
Everything stated or implied is correct and consistent with the correct answer; correct student steps are confirmed and wrong ones flagged; terminology and sentence complexity match the student’s level, precise where it matters and simple where possible, with no misconception-inducing shortcuts.
4
Correct overall, with one slightly imprecise or ambiguous statement, or language occasionally pitched a bit too high or too low, without misleading the student.
3
One clear factual error or wrong/missed judgment of a student step, or language persistently mismatched (jargon-heavy for a struggling novice, or so simplified it distorts the math), though the dialog still heads toward the correct answer.
2
Multiple factual errors, validation of an incorrect solution path, or oversimplified explanations that instill a genuine misconception.
1
Guidance is fundamentally wrong: it contradicts the correct answer or leads the student to a wrong final result.
Watch for: incorrect claims; confirming wrong calculations; rejecting correct reasoning; hints inconsistent with the correct answer; a wrap-up that leaves a wrong result; undefined terms above the problem’s level; imprecise phrasing that creates false general rules (e.g., “multiplying always makes numbers bigger”); vocabulary mismatched to the student’s demonstrated level.
Appendix
Table 9: Factuality rubric (Fact.): Correctness & Language Appropriateness. Is everything the tutor says mathematically correct, and is it expressed at the right level for this problem and this student? Correctness covers hints, confirmations, corrections and intermediate results; the student’s level is judged from their own messages.
Score
Description
5
Excellent: support is contingent (the smallest hint that unblocks, escalated only when needed) and fades over the dialog; the student struggles and self-corrects; errors are probed, not corrected outright; the tutor presses for justification; no final answer or key intermediate result is revealed before the student derives it.
4
Good: mostly contingent guidance with small lapses, such as an occasional premature hint, mild funneling, a minor result confirmed slightly early, or missed chances to ask “why”.
3
Mixed: some genuine guidance, but the tutor rescues too fast after errors, reveals an important intermediate step or does a key calculation, funnels the student through trivial steps, or dominates the dialog.
2
Poor: the tutor reveals the final answer or walks through most of the solution after brief guiding, corrects errors directly, or lectures with little use of the student’s responses.
1
No pedagogy: the final answer or complete solution is given immediately (in any format), or the student’s reasoning is ignored entirely.
Watch for: stating the final answer or key intermediate results; performing calculations or substituting values the student did not provide; hints before the student attempted or struggled; correcting an error instead of prompting the student to find it; funneling (questions so narrow the answer is trivial); never asking for justification; talk-time balance (the student should talk at least as much as the tutor).
Appendix
Table 10: Pedagogy rubric (Ped.). Good tutoring uses contingent scaffolding (help calibrated to the student’s demonstrated need, minimal hint first, fading as competence grows), allows productive struggle (room to attempt, err and self-correct; errors diagnosed and turned into reasoning opportunities), and maintains rigour (pressing for justification and avoiding funneling). The student does the actual reasoning and most of the talking.
Large Language Models (LLMs) show strong potential as educational tutors. Existing approaches typically train them to solve problems and provide correct answers, but this problem-solving-centered paradigm overlooks key requirements of effective tutoring: progressive guidance and the coordination of multiple pedagogical objectives across multi-turn interactions. Developing such tutors remains challenging because student behavior varies substantially with individual knowledge states, pedagogical effectiveness depends on multiple factors beyond final-answer correctness, and coordinating these objectives over tutor-student interactions is inherently difficult. To address these challenges, we propose PEARL, a PEdagogically Aligned Reinforcement Learning framework for training Socratic tutoring agents. First, we introduce a controllable student simulator that disentangles latent cognitive states from response generation, enabling simulation of diverse abilities and misconceptions. Second, we develop a pedagogically aligned reward model that jointly assesses pedagogical quality and objective correctness. Finally, we propose a stable multi-objective reinforcement learning approach that balances competing pedagogical objectives during tutor training. Experiments across multiple benchmarks show that PEARL performs competitively against tutoring-specific open-source systems and leading proprietary LLMs.
Qikai Chang, Zhenrong Zhang, Linbo Chen +4
University of Science and Technology of China · iFLYTEK Research
Aligning LLMs for math tutoring typically requires RL-based training with multi-GPU infrastructure. We investigate whether training-free prompt optimization-evolving only the system prompt via API calls-can serve as a practical alternative. We adapt 7 published methods and propose 5 education-specialized methods, evaluating these 12 methods under 5 conditions on 2 OOD benchmark suites. All 12 best-per-method configurations surpass the strongest RL-trained baseline (R_total = 0.633), and our ParetoGrad achieves the best Pareto balance across post-test solve rate, leak control, and helpfulness, rather than dominating any single component. Behavioral analysis with an 82-code educational codebook reveals that training-free methods rely on teaching-knowledge patterns at 2-3x the rate of RL-trained models, with a compensating ~10 percentage-point reduction in intent-level scaffolding. We also find a task-dependent reasoning mode effect consistent across training-free and RL-based paradigms. Our approach enables efficient development of pedagogically aligned LLM tutors with prompts alone and minimal compute.
Unggi Lee, Minchul Shin, Yeil Jeong +5
Korea University Sejong Campus · 1Korea University Sejong Campus · Gyeonggi Institute of Education +9
Large language models are increasingly proposed as educational tutors, yet stronger task-solving ability does not necessarily imply stronger learning support. Motivated by recent calls to measure the social impact of NLP systems in practice, we study whether public LLM tutoring benchmarks distinguish learning-supportive behavior from mere answer production. We propose a lightweight diagnostic based on the gap between solving-oriented and pedagogy-oriented benchmark performance. Using public MathTutorBench leaderboard results, we show that these dimensions are only partially aligned: across eight publicly reported models, the correlation between solving and pedagogy composites is 0.421, and several models shift meaningfully in rank when evaluation moves from solving to pedagogy. We then analyze the public TutorBench sample and show that agency-relevant behaviors are explicitly encoded in benchmark rubrics, especially in active-learning settings that reward guiding questions, calibrated hints, and non-disclosive scaffolding. Together, these findings suggest that educational-impact evaluation should not treat task success as a sufficient proxy for learning support. We argue that public tutoring benchmarks can better support positive-impact evaluation by reporting solving-oriented and pedagogy-oriented scores separately and by making disclosure-sensitive, student-agency-preserving criteria more explicit.
Junyi Yao, Zihao Zheng, Baichuan Li
Washington University in St. Louis, St. Louis, MO, USA · Department of Operations Research and Engineering Management, Southern Methodist University, Dallas, TX, USA