Corrective feedback is among the best-evidenced drivers of second-language acquisition, yet corrections delivered during lessons rarely accumulate into an actionable view of grammar mastery. Prompted frontier models can provide such a view from learner--tutor lesson transcripts, but they are costly at scale. We close this gap by fine-tuning Qwen3.5 small language models (SLMs) on filtered and rebalanced teacher-generated supervision, then deploying an efficient 0.8B model in an end-to-end grammar mastery tracker for all English learners on our platform. Internalizing the annotation contract into adapter weights enables pairing the 0.8B model with a compact matched prompt rather than verbose instructions. On two human-curated benchmarks, both the deployed 0.8B model and a 4B reference comparator outperform prompted GPT-5.4 and GPT-5.6 Sol in precision and recall under nested matching criteria of increasing strictness: concept, evidence span, and correctness. The deployed 0.8B SLM reduces serving cost by approximately 16×. A feature-level online experiment shows significant gains in learner engagement (+15.8%) and key business metrics, including scheduled hours (+2.1%) and GMV from new lessons (+13.2%).
Figures & tables
Figure 1: Product view of the aggregated grammar mastery score (left) and per-concept breakdown (right; teal values indicate recent mastery gains).
Statistic
Golden
Golden2
Utterances
689
2,977
Grammar concepts
34
34
Concept annotations
3,598
17,400
Annotations / utterance
5.22
5.84
Annotations marked incorrect
592 (16.5%)
6,075 (34.9%)
Utterances with ≥ 1 error
345 (50.1%)
2,845 (95.6%)
Table 1: Composition of the two golden evaluation benchmarks under the common 34-concept evaluation scope. Annotation counts are computed after scope filtering.
GPT-5.4
GPT-5.6 Sol
Qwen3.5-4B
Qwen3.5-0.8B
Prompted (prod.)
Prompted
Base
Fine-tuned
Base
Fine-tuned †
Criterion
P (%)
R (%)
P (%)
R (%)
P (%)
R (%)
P (%)
R (%)
P (%)
R (%)
P (%)
R (%)
Golden
L1: concept
82.34
76.35
87.69
89.66
6.12
33.69
91.26
93.75
0.20
4.53
91.68
94.36
L2: + span
73.20
67.87
79.59
81.38
3.72
20.46
85.82
88.16
0.02
0.47
86.88
89.41
L3: + correctness
67.81
62.87
74.37
76.04
3.01
16.56
81.28
83.49
0.01
0.19
82.74
85.16
Table 2: Micro precision and recall (%) on Golden and Golden2. GPT-5.4 is the production prompted model; GPT-5.6 Sol is a quality-only comparator. Bold marks the best value in each row. Base denotes untuned Qwen3.5 models; fine-tuned denotes selected full-pipeline checkpoints (raw-supervision controls in Table 3 ). The two fine-tuned models use independently tuned recipes; the 0.8B checkpoint, marked † , combines a compact prompt with separately tuned hyperparameters, so differences between sizes are not a controlled comparison. Bootstrap 95% confidence intervals appear in Appendix Table 7 ; key-macro results in Appendix Table 8 . All fine-tuned Qwen–GPT-5.4 differences are significant under paired randomization tests ( p<0.0001 ).
Model
Training data
L3 P
L3 R
Mean-34 L3 R
Rare-9 L3 R
Incorrect L3 R
Golden2
Qwen3.5-4B
Raw supervision
52.47
41.98
34.83
22.04
29.35
Full pipeline
80.84
78.89
69.44
67.35
67.56
Qwen3.5-0.8B
Raw supervision
52.62
40.93
35.25
21.05
29.68
Full pipeline
82.55
80.88
72.64
66.86
69.37
Golden
Table 3: Pipeline ablation (%) on Golden2 and Golden. Bold marks the best value in each column within each benchmark. Mean-34 is equal-weight recall over all concepts; Rare-9 is equal-weight recall over the nine concepts fixed from the raw training split; Incorrect restricts recall to gold incorrect-use annotations. Full-pipeline rows are the selected checkpoints in Table 2 .
Metric
Effect
p
Engagement with their progress
+15.80%
<0.001
Scheduled hours (CUPED)
+2.06%
0.021
Confirmed hours from new tutoring
+10.87%
0.020
GMV from new lessons
+13.20%
0.039
Table 4: Relative effects in the online experiment at the fixed pre-scaling evaluation snapshot (all corresponding 95% CIs exclude zero; bold marks p<0.05 ).
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Qwen3.5 4B
Qwen3.5 0.8B
Qwen3 0.6B
Prompt design
Detailed
Compact
Compact
Max. sequence length
7,000
2,500
2,500
LoRA r/α
16 / 32
128 / 64
128 / 64
LoRA dropout
—
0.05
0.05
LoRA targets
—
All linear
All linear
Learning rate
2×10−5
1×10−4
1×10−4
Appendix
Table 5: Recorded fine-tuning configurations for Qwen3.5-4B, Qwen3.5-0.8B, and Qwen3-0.6B. Effective batch size (256 for the compact runs) is the product of per-device batch size (1), gradient accumulation steps (32), and 8 distributed GPU workers ( 1×32×8=256 ); the 4B run used 1 worker with accumulation 16 ( 1×16×1=16 ). Dashes indicate settings not retained in the available run record.
Model
Val. loss
Epoch
Parse errors
Qwen3.5-4B †
0.2201
15.00
0 / 2,977 (0.00%)
Qwen3.5-0.8B
0.2245
7.50
1 / 2,977 (0.03%)
Qwen3-0.6B
0.2885
8.70
1,710 / 2,977 (57.44%)
Appendix
Table 6: Minimum teacher-forced loss on each run’s training-validation split and structured-output parse errors on Golden2. Bold highlights format adherence ( ≤0.03% parse errors) and the lowest validation loss within each distinct recipe (0.2201 for detailed 4B; 0.2245 for compact 0.8B). The 0.8B and 0.6B runs shared prompt and optimization hyperparameters but selected different early-stopping checkpoints. † The 4B checkpoint used a different prompt, training target, and schedule; its loss is descriptive rather than a controlled cross-model comparison.
Criterion
GPT-5.4 P / R (%)
Qwen3.5-4B FT P / R (%)
Qwen3.5-0.8B FT P / R (%)
Golden (selected checkpoints)
Macro L1: concept
74.08 / 75.34
84.21 / 91.43
85.60 / 89.82
Macro L2: + span
52.44 / 51.65
72.62 / 72.93
75.47 / 75.70
Macro L3: + correctness
45.77 / 45.08
64.48 / 64.82
67.84 / 68.07
Golden2 (selected checkpoints)
Macro L1: concept
72.26 / 68.31
84.41 / 84.85
86.29 / 87.65
Appendix
Table 8: Key-macro precision and recall (%) assigning equal weight to each distinct target matching key at each level, corresponding to the micro results in Table 2 . Bold marks the best value for each metric in each row.
Concept
Raw train
Golden2
Present perfect continuous
187
5
Causative / permissive
259
12
Past perfect
334
9
Zero conditional
335
13
Second conditional
490
17
Reported speech
574
19
Appendix
Table 9: Pre-fixed Rare-9 concepts and annotation support. Raw-train counts determined subset membership; Golden2 counts are reported only as evaluation support.
Figure 3: Simplified production g6.2xlarge native-vLLM topology. The application backend and internal inference API are collapsed into the inference-client block. SageMaker distributes calls across independently provisioned one-GPU containers.
Phase
RPS
Requests
Errors
p50 (s)
p99 (s)
Instances
p1
0.9
1,070
0
4.43
6.87
0→1
p2
1.6
1,955
0
4.71
7.66
1→2
p3
3.2
3,817
2
4.53
20.07
2→5
p4
6.4
7,738
0
4.90
7.64
5→5
p5
1.9
2,246
0
4.23
6.46
5→5
p6
1.0
1,168
0
4.06
6.25
5→4
Appendix
Table 10: Predeployment autoscaling test. “Instances” gives the actual transition during each traffic phase; the initial zero denotes endpoint startup before the configured one-instance minimum became active. Latency is measured per request.
Field
Type
Definition
Original
Text
Learner utterance extracted from the ASR lesson transcript.
Corrected
Text
Upstream LLM correction preserving the intended meaning.
Concept
Enum
One of 34 target grammar concepts.
Span
Text
Exact contiguous substring of the original utterance.
Correct
Boolean
Whether the concept is used correctly in that span.
Explanation
Text
Corrective feedback when usage is incorrect; empty otherwise.
Appendix
Table 11: Logical input and output schema for grammar concept annotation. The first two fields form the input; the remaining fields describe each output annotation.
Metric
Golden2
Synthetic
L1 concept micro F1
92.58
93.50
Partial-span F1
94.53
93.75
Exact-span F1
87.61
62.92
Parse errors
1 / 2,977
0 / 200
Appendix
Table 12: Performance of the same Qwen3.5-0.8B checkpoint under real and synthetic evaluation distributions. Scores are percentages. Span axes evaluate localization independently of concept classification; exact-span performance is diagnostic because span-boundary conventions were not calibrated across annotators.