Adapting general-purpose large language models to specific tasks requires substantial human effort in designing data and training strategies. Sustaining improvement is especially challenging because model updates change the error distribution, requiring strategies to be continually refined. We introduce ImproveAnyTask, an autonomous post-training harness that improves task performance under a limited compute budget. Drawing inspiration from gradient-based parameter optimization, the harness organizes adaptation into error attribution, update-direction selection, and executable model updates. It combines metric-level and case-level analysis to identify a focal problem, then investigates research-backed strategies and compares their reported gains and reproduction difficulty. The selected strategy is translated into training data and a training configuration, with small-scale execution checks preceding full post-training. Subsequent evaluation guides model selection and further adaptation, while validated strategies and scripts are retained for reuse. Across 11 tasks, ImproveAnyTask achieves mean gains of 18.29 and 11.97 percentage points on the Base and Instruct models, respectively, with a maximum gain of 41.96 points, under a 24-hour budget with resources equivalent to eight H20 GPUs.
Figures & tables
Figure 1: Task-specific improvement of Qwen3.5-4B-Base. Scores across 11 tasks in five domains under resources equivalent to eight H20 GPUs for up to 24 hours or three rounds. ImproveAnyTask yields a mean gain of 18.29 points and a maximum of 41.96 points on LiveCodeBench v6. BFCL excludes Web Search and renormalizes the remaining weights.
Figure 2: Overview of ImproveAnyTask. The harness treats task adaptation as iterative optimization: evaluation and dual-level error attribution identify a focal problem; research-grounded strategy exploration selects an update direction; data preparation and post-training produce a candidate model; and re-evaluation determines the next checkpoint and records reusable assets.
Qwen3.5-4B-Base
Qwen3.5-4B-Instruct
Benchmark
Original
Codex
Claude Code
Ours
Original
Codex
Claude Code
Ours
Knowledge & STEM
MMLU-Pro
52.41
53.53 ( ↑ 1.12)
52.27 ( ↓ 0.13)
59.85 ( ↑ 7.44)
58.66
53.41 ( ↓ 5.25)
57.27 ( ↓ 1.38)
69.33 ( ↑ 10.66)
SuperGPQA
32.10
33.97 ( ↑ 1.86)
38.79 ( ↑ 6.68)
47.58 ( ↑ 15.48)
49.78
49.54 ( ↓ 0.23)
47.04 ( ↓ 2.73)
47.77 ( ↓ 2.00)
GPQA-Diamond
47.02
45.23 ( ↓ 1.78)
51.78 ( ↑ 4.76)
69.64 ( ↑ 22.61)
75.59
73.21 ( ↓ 2.38)
78.57 ( ↑ 2.97)
80.35 ( ↑ 4.76)
Mathematical & Verifiable Reasoning
Table 1: Mean results over two independent optimization runs across 11 tasks. Each run selects its post-trained candidate by validation score, excluding Original. BFCL uses the renormalized no-Web metric. All values are percentages.
Figure 3: IFEval optimization over 100 hours. The dashed line marks 24 hours. Checkpoints are equally spaced; changes compare adjacent displayed checkpoints. C5 resumes from C3.
Benchmark
Original
w/o EA
w/o UD
w/o ER
Full Harness
MMLU-Pro
58.66
61.20 ( ↑ 2.53)
60.30 ( ↑ 1.63)
52.51 ( ↓ 6.15)
69.33 ( ↑ 10.66)
MATH-500
51.29
56.00 ( ↑ 4.70)
56.70 ( ↑ 5.41)
54.11 ( ↑ 2.82)
63.52 ( ↑ 12.23)
LiveCodeBench v6
11.49
24.77 ( ↑ 13.28)
33.03 ( ↑ 21.54)
28.57 ( ↑ 17.07)
40.62 ( ↑ 29.12)
TAU2-Bench
36.68
39.21 ( ↑ 2.52)
41.36 ( ↑ 4.68)
38.09 ( ↑ 1.41)
44.94 ( ↑ 8.25)
IFEval
82.78
82.35 ( ↓ 0.43)
83.22 ( ↑ 0.43)
80.17 ( ↓ 2.61)
83.66 ( ↑ 0.87)
Overall Average
48.18
52.70 ( ↑ 4.52)
54.92 ( ↑ 6.74)
50.69 ( ↑ 2.50)
60.41 ( ↑ 12.23)
Table 2: Ablation results for Qwen3.5-4B-Instruct, averaged over two independent optimization runs. Each component is replaced with a basic alternative while preserving the complete optimization loop. All values are percentages.
Figure 4: Optimization trajectories for Qwen3.5-4B-Instruct on MMLU-Pro and MATH-500. Initial and final scores match Table 2 . Curves connect measured checkpoints for visualization.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Validation
Test
Primary metric
MMLU-Pro
1,805
10,227
Accuracy
SuperGPQA
3,980
22,549
Accuracy
GPQA-Diamond
30
168
Accuracy
MATH-500
75
425
Accuracy
AIME 2025
5
25
Accuracy
IFEval
82
459
Strict prompt-level accuracy
Appendix
Table 3: Validation and test sizes used in all comparisons. TAU2 is split by domain and BFCL by category.
Ckpt.
Parent
Strategy and data construction
Examples
C1
Initial Base
Persona-conditioned constraint tuning. Use released instruction–response pairs covering keyword, formatting, length, and capitalization constraints. Preserve message structure, validate formatting, deduplicate, and audit exact overlap with evaluation queries.
29,900
C2
C1
Rubric-guided response tuning. Combine persona-conditioned examples with approximately 26K released responses selected by best-of-six rubric scoring. Convert the selected pairs to the training message format before merging and filtering.
55,998
C3
C2
Multi-source instruction mixture tuning. Combine approximately 30K persona-conditioned, 24K rubric-selected, and 16K Nemotron examples. Retain non-thinking responses from the additional source, remove visual markers, and filter by length.
69,900
C4
C3
Execution-verified targeted constraint tuning. Mix approximately 23.4K verified synthetic examples with 13K replay examples. Target character and keyword frequencies, punctuation, capitalization, and sentence counts identified by the preceding error analysis.
36,299
C5
C3
Replay-preserving additive constraint tuning. Retain the full 69,900-example C3 mixture and add approximately 23,750 execution-verified examples, then deduplicate and filter. Unlike C4, this trial adds targeted supervision without replacing the broader mixture.
93,000
C6
C5
Residual-error-guided constraint augmentation. Extend the 93K-example mixture with 14,524 verified examples covering remaining lexical, counting, formatting, length, and punctuation errors.
107,000
Appendix
Table 4: Data preparation and training ancestry for the extended IFEval Base experiment. Counts in the last column are final exported examples.
Checkpoints
Learning rate
Sequence limit
C1–C3
1×10−5
4,096
C4
5×10−6
4,096
C5–C6
1×10−5
4,096
C7–C8
1×10−5
8,192
Appendix
Table 5: Training settings for the extended IFEval run. The sequence limit is measured in tokens.
Variant
Procedure omitted
Basic alternative retained
w/o EA
Joint metric- and case-level attribution with explicit focal-problem selection.
An unstructured summary of the evaluation results guides the next action.
w/o UD
Required candidate comparison, reproduction-difficulty assessment, resource verification, and outcome-aware reuse.
Ordinary research followed by adoption of a feasible method. The agent may still inspect resources or reconsider its choice.
w/o ER
Required sample-level checks, short training runs, and the execution-guided refinement procedure before full execution.
Ordinary execution and debugging within the same model-update stage. Repairs remain possible after problems are noticed.
Appendix
Table 6: Basic alternatives used in the ablations. ER denotes Execution-Guided Refinement within Executable Model Updates.
Variant
Step 1
Step 2
Step 3
MMLU-Pro
Full
Train on 20K long mathematical reasoning examples.
Train on 20K concise final-answer examples; repair end-of-sequence handling and premature termination.
Introduce cross-domain supervision using complete, verified teacher responses.
w/o EA
Train on 10K auxiliary knowledge questions with short rationales.
Expand to 18K examples and add output-format anchors.
Directly train on long mathematical reasoning traces.
Switch to final-answer supervision while termination and template issues remain unresolved.
Train on 215K examples without checking teacher-response completeness.
MATH-500
Appendix
Table 7: Recorded strategy and execution paths for the five Instruct-model ablation variants. These records provide context for the comparisons in Table 2 .