Grammar Concept Annotation at Scale: Deployed Fine-Tuned Small Language Models Outperform Prompted Frontier Models
Organizations: Preply
Abstract
Corrective feedback is among the best-evidenced drivers of second-language acquisition, yet corrections delivered during lessons rarely accumulate into an actionable view of grammar mastery. Prompted frontier models can provide such a view from learner--tutor lesson transcripts, but they are costly at scale. We close this gap by fine-tuning Qwen3.5 small language models (SLMs) on filtered and rebalanced teacher-generated supervision, then deploying an efficient 0.8B model in an end-to-end grammar mastery tracker for all English learners on our platform. Internalizing the annotation contract into adapter weights enables pairing the 0.8B model with a compact matched prompt rather than verbose instructions. On two human-curated benchmarks, both the deployed 0.8B model and a 4B reference comparator outperform prompted GPT-5.4 and GPT-5.6 Sol in precision and recall under nested matching criteria of increasing strictness: concept, evidence span, and correctness. The deployed 0.8B SLM reduces serving cost by approximately 16. A feature-level online experiment shows significant gains in learner engagement () and key business metrics, including scheduled hours () and GMV from new lessons ().
Figures & tables
| Statistic | Golden | Golden2 |
|---|---|---|
| Utterances | 689 | 2,977 |
| Grammar concepts | 34 | 34 |
| Concept annotations | 3,598 | 17,400 |
| Annotations / utterance | 5.22 | 5.84 |
| Annotations marked incorrect | 592 (16.5%) | 6,075 (34.9%) |
| Utterances with 1 error | 345 (50.1%) | 2,845 (95.6%) |
| GPT-5.4 | GPT-5.6 Sol | Qwen3.5-4B | Qwen3.5-0.8B | |||||||||
| Prompted (prod.) | Prompted | Base | Fine-tuned | Base | Fine-tuned † | |||||||
| Criterion | P (%) | R (%) | P (%) | R (%) | P (%) | R (%) | P (%) | R (%) | P (%) | R (%) | P (%) | R (%) |
| Golden | ||||||||||||
| L1: concept | 82.34 | 76.35 | 87.69 | 89.66 | 6.12 | 33.69 | 91.26 | 93.75 | 0.20 | 4.53 | 91.68 | 94.36 |
| L2: + span | 73.20 | 67.87 | 79.59 | 81.38 | 3.72 | 20.46 | 85.82 | 88.16 | 0.02 | 0.47 | 86.88 | 89.41 |
| L3: + correctness | 67.81 | 62.87 | 74.37 | 76.04 | 3.01 | 16.56 | 81.28 | 83.49 | 0.01 | 0.19 | 82.74 | 85.16 |
| Model | Training data | L3 P | L3 R | Mean-34 L3 R | Rare-9 L3 R | Incorrect L3 R |
|---|---|---|---|---|---|---|
| Golden2 | ||||||
| Qwen3.5-4B | Raw supervision | 52.47 | 41.98 | 34.83 | 22.04 | 29.35 |
| Full pipeline | 80.84 | 78.89 | 69.44 | 67.35 | 67.56 | |
| Qwen3.5-0.8B | Raw supervision | 52.62 | 40.93 | 35.25 | 21.05 | 29.68 |
| Full pipeline | 82.55 | 80.88 | 72.64 | 66.86 | 69.37 | |
| Golden | ||||||
| Metric | Effect | |
|---|---|---|
| Engagement with their progress | +15.80% | |
| Scheduled hours (CUPED) | +2.06% | |
| Confirmed hours from new tutoring | +10.87% | |
| GMV from new lessons | +13.20% |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Qwen3.5 4B | Qwen3.5 0.8B | Qwen3 0.6B |
|---|---|---|---|
| Prompt design | Detailed | Compact | Compact |
| Max. sequence length | 7,000 | 2,500 | 2,500 |
| LoRA | 16 / 32 | 128 / 64 | 128 / 64 |
| LoRA dropout | — | 0.05 | 0.05 |
| LoRA targets | — | All linear | All linear |
| Learning rate |
| Model | Val. loss | Epoch | Parse errors |
|---|---|---|---|
| Qwen3.5-4B † | 0.2201 | 15.00 | 0 / 2,977 (0.00%) |
| Qwen3.5-0.8B | 0.2245 | 7.50 | 1 / 2,977 (0.03%) |
| Qwen3-0.6B | 0.2885 | 8.70 | 1,710 / 2,977 (57.44%) |
| Criterion | GPT-5.4 P / R (%) | Qwen3.5-4B FT P / R (%) | Qwen3.5-0.8B FT P / R (%) |
| Golden (selected checkpoints) | |||
| Macro L1: concept | 74.08 / 75.34 | 84.21 / 91.43 | 85.60 / 89.82 |
| Macro L2: + span | 52.44 / 51.65 | 72.62 / 72.93 | 75.47 / 75.70 |
| Macro L3: + correctness | 45.77 / 45.08 | 64.48 / 64.82 | 67.84 / 68.07 |
| Golden2 (selected checkpoints) | |||
| Macro L1: concept | 72.26 / 68.31 | 84.41 / 84.85 | 86.29 / 87.65 |
| Concept | Raw train | Golden2 |
|---|---|---|
| Present perfect continuous | 187 | 5 |
| Causative / permissive | 259 | 12 |
| Past perfect | 334 | 9 |
| Zero conditional | 335 | 13 |
| Second conditional | 490 | 17 |
| Reported speech | 574 | 19 |
| Phase | RPS | Requests | Errors | p50 (s) | p99 (s) | Instances |
|---|---|---|---|---|---|---|
| p1 | 0.9 | 1,070 | 0 | 4.43 | 6.87 | |
| p2 | 1.6 | 1,955 | 0 | 4.71 | 7.66 | |
| p3 | 3.2 | 3,817 | 2 | 4.53 | 20.07 | |
| p4 | 6.4 | 7,738 | 0 | 4.90 | 7.64 | |
| p5 | 1.9 | 2,246 | 0 | 4.23 | 6.46 | |
| p6 | 1.0 | 1,168 | 0 | 4.06 | 6.25 |
| Field | Type | Definition |
|---|---|---|
| Original | Text | Learner utterance extracted from the ASR lesson transcript. |
| Corrected | Text | Upstream LLM correction preserving the intended meaning. |
| Concept | Enum | One of 34 target grammar concepts. |
| Span | Text | Exact contiguous substring of the original utterance. |
| Correct | Boolean | Whether the concept is used correctly in that span. |
| Explanation | Text | Corrective feedback when usage is incorrect; empty otherwise. |
| Metric | Golden2 | Synthetic |
|---|---|---|
| L1 concept micro F1 | 92.58 | 93.50 |
| Partial-span F1 | 94.53 | 93.75 |
| Exact-span F1 | 87.61 | 62.92 |
| Parse errors | 1 / 2,977 | 0 / 200 |