cs.CLOct 5, 2026

The Assistance Dilemma: Learning to Teach via Multi-Turn Reinforcement Learning

Authors: Jakub Macina, Manu Kapur, Mrinmaya Sachan

Organizations: ETH Zurich

Abstract

Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors reward the student's success on the tutored problem with the tutor's words still in context. The reward is then easiest to raise by telling the student the answer, and a tuned penalty is needed to reduce telling. Drawing on learning sciences, we introduce a masked near-transfer post-test: the student is tested on an unseen variant of the tutored problem with the tutor's utterances masked, so the reward can rise only through what the student wrote in its own turns. This discourages cognitive offloading by the student and allows the continuous penalty to be replaced by two binary reward gates (factual correctness of tutor response, no solution handover). A leave-one-out ablation shows that the learning-gain reward on its own does not separate teaching from telling: the gates reduce solution handover while the near-transfer post-test improves out-of-domain transfer. Using these reward designs we develop Eduardo, a multi-turn RL recipe for training LLM tutors, and use it to train 4B, 9B, 14B and 27B models from two distinct LLM architectures. Our post-trained Eduardo-27B model matches Gemini-3.1-Pro on MathTutorBench and Claude Opus 4.8 on TutorMoments at 2.4-6.2x fewer thinking tokens than frontier models, which matters for interactive tutoring. Without being named in the reward, the model more than doubles its use of the push-for-justification teacher move while support fading (e.g., assigning independent work), whose payoff lies beyond a single-problem dialog episode, is trained out. We open-source our training environment, an 8,671-problem near-transfer dataset, and trained models for further development.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning

    May 28, 2026Qikai Chang, Zhenrong Zhang, Linbo Chen +4Socratic DialogueOffline Reinforcement Learning

  2. LLMs Are Already Good Tutors: Training-Free Prompt Optimization for Pedagogical Math Tutoring

    May 26, 2026Unggi Lee, Minchul Shin, Yeil Jeong +5TutorsTraining-Free

  3. Beyond Helpfulness: A Teaching-over-Solving Diagnostic for Measuring Educational Impact in LLM Tutors

    Jun 15, 2026Junyi Yao, Zihao Zheng, Baichuan LiTutorsComputerized Adaptive Testing