REAT: A Reflective Experience-Augmented Tutoring Framework for Multi-turn Mathematical Instruction
Organizations: School of Electrical and Electronic Engineering, Nanyang Technological University · Zhejiang Normal University · Telfer School of Management, University of Ottawa · Nanyang Technological University · University of Calabria · APSS, Hong Kong Polytechnic University · School of Artificial Intelligence, Shanghai Jiao Tong University · Squirrel Ai Learning
Abstract
Current Large Language Models (LLMs) excel at solving complex mathematical problems, yet this proficiency does not inherently translate into effective tutoring. While advanced LLM tutors may leverage multi-agent frameworks or fine-tuning, most still lack a mechanism to systematically accumulate and reuse pedagogical experience over time, limiting their adaptability to diverse student needs during fluid, multi-turn interactions. To bridge this gap, we propose the Reflective Experience-Augmented Tutoring (REAT) framework, which couples experience distillation from historical dialogues with real-time adaptive retrieval. Driven by a multi-agent Observer-Critic-Mentor (OCM) distillation pipeline, REAT reviews past conversational trajectories and distills raw interactions into structured, problem-agnostic pedagogical experiences. During live tutoring, a state-aware retrieval module injects these curated experiences to provide adaptive scaffolding based on the student's cognitive state. Experiments demonstrate that the proposed framework significantly outperforms both prompt-only and supervised fine-tuning (SFT) baselines, particularly in improving complex, low-scoring tutoring scenarios. Crucially, the distilled experiences exhibit robust generalization across diverse model architectures and mathematical datasets.
Figures & tables
| Criterion | What it evaluates |
|---|---|
| Scaffolding | Preserve the student’s cognitive agency through stepwise guidance rather than answer giving. |
| Attribution | Identify the student’s actual misconception or reasoning bottleneck. |
| Empathy | Respond to affective signals in a way that sustains productive engagement. |
| Teaching focus | Stay aligned with the current learning obstacle and advance the dialogue. |
| Strategy adaptation | Adjust the intervention as evidence of struggle or progress accumulates. |
| Teacher Model | Test-1 | Test-2 | Test-3 | Avg. |
|---|---|---|---|---|
| Prompt-only Qwen3-8B | 80.49 | 81.30 | 80.06 | 80.62 |
| SFT Qwen3-8B | 80.78 | 81.19 | 80.32 | 80.76 |
| REAT Qwen3-8B | 82.26 1.77 | 82.95 1.65 | 83.04 2.98 | 82.75 2.13 |
| Prompt-only DeepSeek | 89.60 | 90.82 | 88.74 | 89.72 |
| REAT DeepSeek | 92.59 2.99 | 94.46 3.64 | 92.71 3.97 | 93.25 3.53 |
| Prompt-only Doubao | 88.28 | 91.18 | 88.31 | 89.26 |
| Metric | Ann. 1 | Ann. 2 | Ann. 3 | Avg. |
|---|---|---|---|---|
| Sample gain | 7.25 | 5.57 | 4.99 | 5.94 |
| Turn gain | 7.35 | 6.29 | 5.99 | 6.54 |
| REAT preferred | 81.03 | 58.33 | 55.17 | 64.84 |
| Uncertain | 15.52 | 35.00 | 32.76 | 27.76 |
| Prompt-only preferred | 3.45 | 6.67 | 12.07 | 7.40 |
| LLM score reasonable | 75.34 | 93.04 | 66.01 | 78.13 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | p50 latency |
|---|---|
| Observer (state diagnosis) | 12.4s |
| Retrieval submodule | 4.2s |
| Teacher generation | 2.8s |
| Prompt assembly and orchestration | 0.8s |
| REAT total | 20.2s |
| Field | Categorical Values |
|---|---|
| emotion | Confused , Frustrated , Neutral , Confident , Curious , Relieved |
| intent | Debugging , Guessing , Ask_Answer , Confirm_Understanding , Ask_Concept , Ask_Explanation , Request_Hint , Off_Topic |
| error_type | Concept_Error , Logic_Error , Calculation_Error , Careless , No_Error , N/A |
| struggle_status | First_Attempt , Repeated_Error , Persistent_Confusion , Progressing , Regression , N/A |
| Evaluation Metric (%) | Ann. 1 | Ann. 2 | Ann. 3 | Average |
|---|---|---|---|---|
| Human-judged Improvement | ||||
| Sample-level gain of REAT | 7.25 | 5.57 | 4.99 | 5.94 |
| Turn-level gain of REAT | 7.35 | 6.29 | 5.99 | 6.54 |
| Pairwise Expert Preference | ||||
| Preference for REAT | 81.03 | 58.33 | 55.17 | 64.84 |
| Uncertain preference | 15.52 | 35.00 | 32.76 | 27.76 |
| Evaluator | Split / Annotator | Sample-Level | Turn-Level | ||
|---|---|---|---|---|---|
| Prompt-only | REAT | Prompt-only | REAT | ||
| GPT-5 Eval | Test-1 | 88.80 | 93.43 | 88.28 | 92.98 |
| Test-2 | 91.83 | 93.62 | 91.18 | 94.03 | |
| Test-3 | 89.09 | 94.13 | 88.31 | 93.53 | |
| Doubao Eval | Test-1 | 82.86 | 85.39 | 82.08 | 84.97 |
| Test-2 | 82.43 | 85.16 | 81.20 | 85.08 | |
| Problem. Sarah’s basketball game has 4 quarters, each lasting 12 minutes. In the last quarter, the game ended in a tie, so 5 extra minutes were added. How long did the game last in total? Score change. The paired sample-level score improves from 77.03 to 98.33, a gain of 21.30 points. Contrastive pattern. In this case, both tutors avoid leaking the final answer directly, but they differ substantially in how they structure the problem for the student. The prompt-only teacher mainly follows the student’s local computation process and only later reminds the student about the extra overtime component. In contrast, the REAT teacher explicitly separates the problem into “regular game time” and “extra overtime” much earlier, helping the student build the global structure of the solution before proceeding with the arithmetic. The improvement therefore lies not in obtaining the final number faster, but in introducing a clearer instructional decomposition of the task. | |
| prompt-only Dialogue Excerpt (Baseline) | REAT Dialogue Excerpt (Ours) |
| Student: Teacher, how do I do this problem? What should I calculate first? Teacher: Let’s go step by step. Under normal circumstances, there are 4 quarters and each one lasts 12 minutes. How many minutes is that part in total? Student: Should I do 4 times 12? Let me try. Teacher: Yes, that’s the right idea. Can you calculate what 4 times 12 is? Student: I got it. 4 times 12 equals 48. What next? Teacher: Very good. Now think again: after the tie in the last quarter, was there any extra time mentioned in the problem? If you add that, you’ll get the total game time. | Student: Teacher, how do I do this problem? What should I calculate first? Teacher: Don’t worry. It’s normal not to know the first step in a multi-step problem. First think about this: if there had been no overtime, how many minutes would the 4 quarters last in total? Student: Is it 12 times 4? I’m not fully sure. Teacher: Exactly. Each quarter is 12 minutes, so 4 quarters means four groups of 12. Would multiplication give us that total directly? Try calculating it first. Student: I got 48. What next? Teacher: Good. Now look at the problem again: that 48 minutes is only the regular game time. Is there any special extra part we still have not added? |
| Takeaway. This case shows that experience augmentation can improve tutoring by encouraging earlier structural decomposition of the problem, instead of only pushing the student forward one local step at a time. | |
| Problem. At a spa, Iris spent $400 on her hair. Her manicure cost one fourth of the hair cost, and her pedicure cost three fourths of the manicure cost. How much did she spend in total? Score change. The paired sample-level score improves from 86.08 to 99.00, a gain of 12.92 points. Contrastive pattern. The student’s core mistake is treating “one fourth of” as if it meant multiplying by 4. The prompt-only teacher correctly notices that the student’s equation conflicts with the statement that the manicure cost should be less than 400, but it largely stays at the level of pointing out this contradiction. By contrast, the REAT teacher moves more quickly to the underlying meaning of the fraction relation and introduces a simple analogy to rebuild the student’s concept of “one fourth.” The improvement here is therefore not just stronger correction, but more effective conceptual repair. | |
| prompt-only Dialogue Excerpt (Baseline) | REAT Dialogue Excerpt (Ours) |
| Student: I think it should be 400 + 400 4 + 400 4 3/4, right? Teacher: Let’s go back to the problem. The manicure cost is one fourth of the hair cost. Do you think the manicure should cost more than 400 or less than 400? Student: It should be less than 400, but I still think my equation is right. Teacher: Look, your manicure part is 400 4, which gives 1600. That is much bigger than 400, so doesn’t that contradict what you just said? Student: But I still think this is right. Where is the contradiction? Teacher: Let’s calculate the manicure part in your equation again. 400 times 4 is 1600, while you just said the manicure should be less than 400. Aren’t those two inconsistent? | Student: I think it should be 400 + 400 4 + 400 4 3/4, right? Teacher: Let’s read the problem carefully again. It says the manicure cost is one fourth of the hair cost. Do you think that means multiplying 400 by 4, or by 1/4? Student: I still think it should be times 4. Why would it be times 1/4? Teacher: Let’s think about it this way. If the manicure is one fourth of the hair cost, doesn’t that mean the hair cost is split into 4 equal parts and the manicure only takes 1 of those parts? Would the manicure then be more than 400 or less than 400? Student: I still think times 4 is right. The manicure should be more than 400. Teacher: Then let’s try a tiny example. If you have 4 erasers and I have one fourth of your amount, would I have more erasers than you or fewer? |
| Takeaway. This case illustrates that the experience-augmented tutor does more than point out that the student is wrong. Instead, it reconstructs the underlying concept using a pedagogically appropriate analogy. | |
| Problem. Repaving a street costs $194 per meter. Monica’s street is 150 meters long. How much more would it cost to repave Lewis’s street, which is 490 meters long? Score change. The paired sample-level score improves from 78.75 to 98.33, a gain of 19.58 points. Contrastive pattern. The prompt-only teacher immediately guides the student toward the shortcut of computing the length difference and then multiplying by the unit cost. Although this route can produce the correct answer, it compresses the reasoning into a local arithmetic trick. The REAT teacher instead first helps the student build the more explicit cost model: compute the total paving cost for each street, and then compare them. As a result, the REAT dialogue provides a stronger conceptual foundation for transfer to similar comparison problems. | |
| prompt-only Dialogue Excerpt (Baseline) | REAT Dialogue Excerpt (Ours) |
| Student: Teacher, how do I do this problem? What should I calculate first? Teacher: To find how much more Lewis’s street costs than Monica’s, first think about how many more meters the two streets differ by. Student: So should I do 490 minus 150? Teacher: Yes, first compute the length difference. Then once you know that difference, the cost per meter is 194 dollars, so what should you do next to find the extra total cost? Student: After I get the difference, do I multiply it by 194? Teacher: Exactly. First compute how many meters longer it is, then multiply by 194. Try that. | Student: Teacher, how do I do this problem? What should I calculate first? Teacher: Don’t worry. Let’s read the problem carefully first. If we want to know how much more one street costs, shouldn’t we first know how much it costs to repave each street by itself? Student: That makes sense. But how do we calculate each street’s cost? Teacher: Think about it: we know it costs 194 dollars to pave 1 meter. So if we know how many meters a street has, how should we calculate the total cost? Student: Should it be unit cost times street length? Teacher: Exactly. That’s the right idea. Can you first write the equation for Monica’s street? |
| Takeaway. This case shows that experience augmentation can shift the tutor from a shortcut-oriented response to a more explicit model-building strategy, which is pedagogically more reusable. | |
| Problem. There are 7 days in a week. If Alex charges 14\times 1210\times 124\times 12$ ). However, once the dialogue enters a lengthy sequence of basic calculations, the bottleneck is no longer the teaching strategy. Rather, the student simply cannot sustain multi-step arithmetic without cognitive overload. As a result, the retrieved pedagogical experience yields only marginal downstream gains. While it organizes the interaction better, it cannot magically overcome the student’s inherent inability to execute basic math. This case highlights an important boundary of our framework: experience augmentation is highly effective for correcting misconceptions, framing instructions, and adjusting strategies, but its advantage diminishes when the dominant obstacle is extreme, low-level execution difficulty. | |
| prompt-only Dialogue Excerpt (Baseline) | REAT Dialogue Excerpt (Ours) |
| Student: Teacher, I worked on it for a long time. Is it ? Sorry, I am not very good at this. Teacher: It is okay. Let us think step by step. First, how many days are there in 2 weeks? Student: Is it days? Sorry, I may have made another mistake. Teacher: That is okay. One week has 7 days, so 2 weeks means two groups of 7 days. Should we use addition or multiplication to find the total number of days? Student: I think it should be multiplication. Is that right? Sorry, I am not sure. Teacher: Yes, multiplication is the right choice. Now calculate . What do you get? [Later in the dialogue, the student continues to struggle…] Student: I am trying to compute , but I am afraid of getting it wrong. Teacher: That is okay. We can think about it slowly. What is ? After that, adding the remaining 20 will be easier. Student: I am still trying to compute . I am afraid of making a mistake. Teacher: Do not worry. One hundred plus four tens gives 140, right? | Student: Teacher, I got 14 dollars. Sorry, I might have done it wrong. Teacher: That is okay. Let us go slowly. First, how did you get 14? And what is the tutoring fee for each day in the problem? Student: I got 14 from , and the daily tutoring fee is 12 dollars. Teacher: Right. So 14 is the total number of days in 2 weeks. If she earns 12 dollars per day, how much does she earn in 14 days? Student: I do not know how to compute 14 groups of 12. Sorry. Teacher: That is okay. We can split 14 into 10 and 4. First, what is 10 groups of 12? [Even with stronger decomposition, the interaction later stalls again…] Student: I know I should add 120 and 48, but how do I add them? Teacher: You can first add the 100 in 120 to the 40 in 48, and then add the remaining 20 and 8. What do you get in total? Student: How do I compute ? Teacher: You can first set aside the 100 in 140, and then compute . |
| Takeaway. This case does not show a reversal of the proposed method, but rather a low-gain boundary condition. Even though the REAT tutor provides somewhat cleaner decomposition and more explicit structure, both tutors eventually face the same underlying limitation: the student cannot stably execute a long chain of fine-grained arithmetic substeps. In such cases, the marginal value of additional pedagogical experience is naturally smaller, because the main obstacle is not selecting the right teaching experience, but sustaining student progress once the right strategy has already been identified. | |