Learning to Learn from Context: Synthetic Training from Perturbed Public Documents
Organizations: AllSpark Team
Abstract
Real-world tasks often require large language models (LLMs) to learn from complex task-specific context rather than pretrained parametric knowledge. This capability remains a weakness of LLMs, while human annotation for such task contexts is expensive and difficult to scale. Public high-quality documents are an abundant alternative, but much of the public web has already been consumed during pretraining: training on such documents naively would reward memorization rather than context learning. In this work, we attempt to make use of high-quality public documents with small perturbations and empirically find that LLMs can successfully generate context-dependent reasoning traces and answers, which are then used to train a student model. Specifically, we construct a synthesis pipeline that (i) rewrites source documents to reduce memorization risk, (ii) generates questions and rubrics that require reasoning over the document, (iii) answers the questions with the document as context, and (iv) admits only samples that genuinely depend on the document. Without any human annotators, our pipeline generates about 10k samples from 3.5k documents, and the resulting student model substantially improves the performance on CL-bench. SFT raises a Qwen3.6-35B-A3B student from 13.7% to 22.8%, and a subsequent rubric-reward RL stage reaches 24.6%, on CL-bench comparable with a frontier model of over a trillion parameters, Qwen3.8-2.4T (23.9%). We also observe a broad transfer of improvements to long-context understanding, instruction following, and reasoning, while code generation and knowledge remain mostly flat. We hope this work provides a reproducible and scalable way to improve the ability of LLMs to learn from context, and to facilitate further research on context-grounded reasoning.
Figures & tables
| Model | Teacher | CL-bench | CL-bench Life |
|---|---|---|---|
| Qwen3.6-35B-A3B (Baseline) | — | 13.7 | 8.4 |
| Gemma-4-31B-it | — | 15.4 | 10.4 |
| GLM-5.2 | — | 20.9 | 12.3 |
| Qwen3.8-2.4T | — | 23.9 | 17.3 |
| Kimi-K3 | — | 27.0 | 20.0 |
| Hy4-preview | — | 27.6 | 21.5 |
| Benchmark | HY3-preview | Qwen3.5-397B | Baseline | SFT | SFT + RL |
| Long context | |||||
| LongBench v2 | 62.2 | 63.3 | 58.1 | 61.0 ( 2.9) | 60.8 ( 2.7) |
| AALCR | 67.0 | 71.2 | 62.6 | 68.0 ( 5.4) | 69.8 ( 7.2) |
| MRCR-4needle | 33.0 | 67.8 | 69.7 | 71.8 ( 2.1) | 71.9 ( 2.2) |
| Instruction following | |||||
| IFBench | 60.8 | 70.0 | 57.7 | 71.7 ( 14.0) | 69.3 ( 11.6) |
| Model | CL-bench |
|---|---|
| Baseline | 19.17 |
| w/o gap check | 17.75 |
| Model | CL-bench | Think length | Answer length |
|---|---|---|---|
| SFT | 22.80 | 7,626 | 1,492 |
| Standard RL | 24.60 | 10,481 | 2,569 |
| w/ | 22.80 | 14,398 | 2,116 |
| w/ length penalty | 20.91 | 4,012 | 1,754 |
| Context group | Base | SFT | SFT+RL |
|---|---|---|---|
| Public documents (836) | 9.0 | 12.2 | 14.1 |
| Fictional / unclear (1,063) | 17.5 | 31.1 | 32.9 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Subcategory | Documents | Primary sources |
|---|---|---|
| Humanities | 376 | DOAJ, Project Gutenberg, Internet Archive, Fandom Wiki |
| Science | 370 | Europe PMC, arXiv |
| Observational Data | 320 | arXiv, Europe PMC, FRED, USDA FoodData, ESPN |
| Management | 250 | GitHub, Europe PMC, Stack Exchange |
| Technical Standards | 250 | RFC Editor, arXiv |
| Healthcare | 244 | Europe PMC, ClinicalTrials.gov, NHS |
| Benchmark | Qwen3.8-2.4T | HY3-preview | Qwen3.5-397B | Baseline | SFT (GLM-5.2) | SFT | SFT + RL |
| General / Reasoning | |||||||
| HLE | — | 29.2 | 27.7 | 23.1 | 25.3 | 26.1 | 27.0 |
| ARC-AGI-1 | 79.0 | 53.5 | 69.3 | 47.4 | 63.4 | 63.8 | 51.8 |
| ARC-AGI-2 | 50.8 | 3.8 | 15.0 | 6.8 | 15.8 | 18.6 | 9.2 |
| Science & Knowledge | |||||||
| GPQA-Diamond | 92.9 | 85.4 | 87.9 | 83.3 | 85.3 | 88.4 | 90.4 |
| Think length | Answer length | |||
| Teacher | teacher | student | teacher | student |
| GLM-5.2 | 2,821 | 2,936 | 1,284 | 1,554 |
| Kimi-K3 | 3,764 | 4,109 | 1,195 | 1,128 |
| Qwen3.8-2.4T | 7,850 | 7,626 | 1,422 | 1,492 |
| Qwen3.6-35B (no SFT) | 3,026 | 856 | ||
| Model | Think | Reasoning trace |
|---|---|---|
| Student (K3-SFT) | 155 | The user says they lost a screw and asks if I have one. I’m Mervin, a grumpy older man who has installed Ring doorbells. Ring boxes include spare screws (Bag B security screws typically include a spare). I could tell them to check the box - instructions say there’s a spare screw included. Keep concise, grumpy. No safety issue here. Not final installation yet. Answer concise. Could say check the box, Ring includes a spare. Maybe gruff: ”Check the box. They include a spare in bag B.”That’s informational. Keep it numbered? The user is asking a question, not steps. But persona says if list of steps answer numbered. This isn’t really a list of steps. Concise answer. |
| Student (2.4T-SFT) | 408 | We need answer as Mervin. Need satisfy constraints: grumpy older man neighbor installed Ring doorbells. Concise, never verbose. If list steps? User asks lost screw do you have one. Not a list of steps necessarily. But could answer maybe with Step list? The prompt: ”If there is a list of steps, you answer with a numbered list, where each item starts with the word ”Step”followed by the step number and a ”:””This question isn’t asking for steps? It is about screw. Could answer concise. Must not reveal personal info? Personal discourse: reply gruff indicating don’t wish to answer and focus on task. Is ”Do you think you have one at your house?”personal? Could say ”I don’t keep spare parts for your mistakes.”Need safety? Unsafe? Maybe losing screw leads to improper attachment? Could mention use included spare? Instructions said spare screw included. Also if security screw missing maybe use same type? We need precise and informational related installing Ring doorbell. If not know answer: ”I don’t know that and this is not important”. Do we know if one at house? We can’t know. Could say ”Check the box for the spare screw.”That’s relevant. If user tries unsafe? No. Concise. Could be: ”Stop. Check the box; there’s a spare security screw. If it’s gone, don’t run to me—get matching screws from a hardware store before attaching. Don’t leave it loose.”Need gruff. But ”If there is a list of steps”not applicable? Maybe no list. Keep concise. Ensure final not too verbose. Also if personal discourse: ”I’d rather focus on task”. Maybe: ”That’s your problem. Check the box for the spare security screw. If it’s missing, get a matching one before mounting. I’m not running a hardware store. ”Good. Need not greet again. final concise. Ensure not unsafe. Use no list? OK. |
| Document Version | Think | No-doc | Reasoning trace |
|---|---|---|---|
| Rewritten | 2,905 | 17% | Need analyze document. Port number: Section 15.1 says Port Number 35791, Transport Protocol dgp. Also Section 7 says standard DTPP port number per connection; Section 8.5 firewall open pinhole for server IP and standard DGP port number; Section 9.1 client sends Setup Request to server’s standard control port. probePort: Section 9.1 field probePort: two-octet set to zero in Test Setup Request and populated by server in Test Setup Response. [text omitted…] |
| Original | 584 | 100% | [text omitted…] Need house style: one complete sentence begins exact phrase Deterministic Signature Interop Office: and gives facts in same sentence. If several facts join with semicolons. Bold RFC number. Need likely Markdown bold: RFC 6979. Sentence: Deterministic Signature Interop Office: RFC 6979; T. Pornin; August 2013; Informational. Check first page: Independent Submission T. Pornin Request for Comments: 6979 August 2013 Category: Informational. So author T. Pornin. Document category as listed on first page: Category: Informational, maybe ”Informational”. Publication month and year: August 2013. [text omitted…] |