Large language models are increasingly expected to support diverse chemical reasoning capabilities within a unified model. One approach is to develop specialized capabilities separately and consolidate them through multi-teacher on-policy distillation, but this raises two questions: how should specialization be organized, and how should specialist guidance be integrated? We introduce ChemOPD, which addresses both. We estimate task affinities from supervised fine-tuning gradients and solve a constrained mixed-integer program(MIP) to construct partially overlapping specialist groups. During distillation, we retain a generalist teacher trained on all tasks so that specialist guidance supplements rather than replaces its supervision. Our anchor-residual objective gradually increases the routed specialist's contribution on student-generated responses. On ChemCoTBench, affinity-guided specialization produces task-dependent gains over the generalist teacher and improves several capabilities beyond semantic task grouping. Yet stronger teacher-side performance does not automatically yield stronger students: with the same specialists and routes, anchor-residual OPD improves most reported metrics over specialist-only distillation and realizes a larger share of the available teacher gains. These results highlight specialization and capability integration as connected but distinct design problems in chemical reasoning.
Figures & tables
Figure 1: Overview of affinity-guided teacher construction and anchor-residual multi-teacher OPD. The MIP converts gradient-based task affinities into overlapping task groups for specialist training. During OPD, the generalist teacher and the fixed-route specialist score the same student-generated prefixes. The generalist’s advantage serves as the anchor, and the specialist residual is introduced gradually to train a unified student.
Figure 2: Performance of teacher configurations across benchmark tasks. Base Student and Base Teacher are the pretrained Qwen3-1.7B and Qwen3-8B models, respectively, while SFT Student is the common all-task SFT initializer. Generalist Teacher is trained on all tasks, whereas Benchmark-Family, Random-Group, and MIP-Routed denote three specialist-routing schemes. Bars show the primary metric for each task. Squares show property change ( Δ ) in panel (b) and Top-1 in panel (c); MechSel is reported by accuracy. Downward arrows mark lower-is-better metrics. Blue annotations summarize the better/tied/worse counts of MIP-Routed relative to Generalist Teacher. Full results are provided in Appendix Table 10 .
Figure 3: Affinity-based task–group compatibility for the selected MIP grouping. Each cell reports the mean compatibility Wij between the column task and the other tasks assigned to the row teacher. Outlined cells indicate the memberships selected by the MIP. LogP and solubility are shared by Teachers 1 and 3.
Editing
Understanding
Method
Add P@1
Delete P@1
Sub. P@1
Murcko Sim.
Equiv. Acc.
FG MAE ↓
Ring MAE ↓
Ring-Sys Sim.
Student, all-task SFT
0.30
0.45
0.38
0.53
0.58
0.11
0.20
0.72
MIP-routed SFT
0.50
0.70
0.62
0.77
0.66
0.08
0.10
0.82
Vanilla OPD
0.35
0.55
0.48
0.55
0.55
0.07
0.15
0.77
Benchmark-family OPD
0.45
0.25
0.40
0.53
0.60
0.10
0.15
0.70
MOPD
0.50
0.20
0.52
0.48
0.61
0.14
0.25
0.77
Table 1: Task-level results across OPD methods. Each panel retains the native metric for each subtask; bold marks the best OPD method in a column before display rounding, and the light-blue row highlights ChemOPD. FG and Ring are lower-is-better; all other metrics are higher-is-better.
Figure 4: Teacher and student gains over the common SFT initializer. For metric i , Hi is the direction-aligned gain of the routed teacher, and Gi is the corresponding student gain; positive values indicate improvement. Panels (a–b) plot Gi against Hi for specialist-only MOPD and ChemOPD across all 31 metrics. Colors denote task families, and the dashed diagonal Gi=Hi marks full recovery of the teacher gain. Panel (c) reports the median ratio Gi/Hi overall and within each task family for the 26 metrics with Hi>0.02 . Gray and blue denote MOPD and ChemOPD, respectively. All gains retain their native metric units; metric columns are descriptive and are not independent trials.
Figure 5: OPD training dynamics. Panels (a–c) show 10-update averages of (a) the student-teacher entropy gap Eb , (b) the signed top-16 supervision signal Ss(16) , and (c) the percentage of token candidates with ∣A^∣>5 . Panel (d) shows the empirical CDF of the pre-clipping actor gradient norm Gs ; the dashed line marks Gs=1 . Annotations compare ChemOPD with MOPD, which uses the same prompts and specialist routes. See Appendix I for definitions.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Family
Subtasks
Primary evaluation dimensions
Editing
Add, Delete, Substitution
P@1 for a valid output satisfying the requested functional-group change
Understanding
Murcko scaffold, Equivalence, functional-group count, ring count, ring-system recognition
Murcko Morgan-Tanimoto; equivalence/ring-system accuracy; FG/ring MAE
Optimization
QED, LogP, solubility, DRD2, JNK3, GSK3B
Property gain Δ and success rate SR(Δ>0) ; validity/scaffold diagnostics
Reaction
Fwd by , Fwd major , MechSel, NEPP, Condition/RCR, Retro
Top-1 chemical match and FTS for SMILES tasks; MechSel accuracy
Appendix
Table 2: 20 chemical reasoning subtasks used in the study. Detailed task semantics and metric calculations are given in Appendix B.2 .
Family
SFT training
Evaluation
Molecular editing
4,497
100
Molecular understanding
7,718
320
Molecular optimization
5,587
600
Reaction reasoning
10,371
575
Total
28,173
1,595
Appendix
Table 3: SFT training instances and evaluation task views by benchmark family.
Subtask
Train
Eval.
Scored
Byproduct prediction
890
100
81
Major-product prediction
2,788
100
100
Mechanism selection
229
100
100
NEPP
1,257
85
85
RCR
3,142
90
90
Retrosynthesis
2,065
100
100
Appendix
Table 4: Reaction-task instance counts under the current protocol. Each released forward-synthesis example contributes one major-product and one byproduct view.
Statistic
Count
Source records
2,165
Normalized product-string groups
1,818
Source reaction classes
87
Singleton product-string groups
1,676
Classes with singleton candidates
82
Training records
2,065
Appendix
Table 5: Construction of the retrosynthesis holdout.
Hyperparameter
Value
Teacher initialization
Qwen3-8B-Base
Student initialization
Qwen3-1.7B-Base
Fine-tuning scope
Full-parameter
Training objective
Assistant-token causal LM loss; prompt masked
Epochs
3
Optimizer
AdamW, β=(0.9,0.999) , weight decay 0.01
Appendix
Table 6: Core supervised fine-tuning settings. Microbatch size, gradient accumulation, and gradient checkpointing were adjusted to model size and GPU allocation; the optimizer and data-processing settings below were shared.
Hyperparameter
Value
Student initialization
All-task SFT Qwen3-1.7B
Training budget
109 distributed optimizer updates
Prompt batch size
256
Rollouts per prompt
4
Actor minibatch
256 prompts / 1,024 rollout sequences globally
Computational microbatching
Dynamic per-GPU packing; 4,096-token budget
Appendix
Table 7: OPD hyperparameters shared by the completed methods. The method-specific batch construction used by Homogeneous-RR and DanceOPD is described in Appendix C.2 ; the implemented distillation signal and actor objective are defined in Appendix I .
Table 8: Affinity-guided soft-overlap groups used by the current reaction protocol experiments. LogP and Solubility occur in both Teacher 1 and Teacher 3; their fixed OPD route is Teacher 3.
Hyperparameter
Value
Probe backbone
Qwen3-8B-Base
Probe examples per task
100 train / 50 validation
Probe batch/sequence length
1 / 4096 tokens
Gradient representation
Last 2 transformer layers; LM head excluded
CountSketch dimension
8192
Probe precision and seed
bfloat16; 42
Appendix
Table 9: Affinity-probe and MIP settings used to construct the three ChemOPD specialists. The selected profile is a practical solver output rather than a claim of a unique globally optimal grouping.
Figure 6: Pairwise compatibility within the three MIP-selected teacher groups. Each panel shows Wij=Sij+Sji for every unique off-diagonal task pair within one group. Positive and negative values indicate above- and below-average compatibility, respectively.
Editing
Understanding
Configuration
Model
Add P@1
Delete P@1
Sub. P@1
Murcko Sim.
Equiv. Acc.
FG MAE ↓
Ring MAE ↓
Ring-Sys Sim.
Student, raw
Qwen3-1.7B
0.15
0.00
0.00
0.59
0.49
0.15
1.05
0.67
Teacher, raw
Qwen3-8B
0.35
0.55
0.08
0.10
0.60
0.16
0.31
0.67
Student, all-task SFT
Qwen3-1.7B
0.30
0.45
0.38
0.53
0.58
0.11
0.20
0.72
Teacher, all-task SFT
Qwen3-8B
0.65
0.70
0.52
0.74
0.70
0.08
0.10
0.82
Benchmark-family SFT
Qwen3-8B
0.75
0.65
0.38
0.64
0.65
0.14
0.15
0.80
Appendix
Table 10: Full task-level results for the teacher constructions. The MIP and benchmark-family profiles use predefined task-to-teacher routes. The shaded row shows the routed profile of the MIP specialists used by ChemOPD. Blank entries indicate tasks not covered by the corresponding human-defined controls.
Figure 7: Direction-aligned teacher gains from Table 10 . The dashed ring is the generalist teacher; outward is better, with MAE signs reversed. Human-control curves remain open for uncovered tasks, while MIP-routed follows its fixed route. Teal labels report pooled pairwise comparisons between MIP-routed and all available human-control results. Each metric contributes one comparison for every human control that covers it. Metrics retain native units and panel-specific scales.
Figure 8: Stability of affinity estimation and MIP grouping under probe resampling. (a) Agreement between each resampled affinity matrix and the reference. Open circles show the ten draws. Light bars show the full range, dark bars show the interquartile range, and diamonds show the median. The top-3 neighbor score is averaged over tasks. (b) Task-level top-3 neighbor Jaccard scores. A value of one means that all three neighbors are unchanged. (c) The objective score of the fixed ChemOPD grouping compared with 10,000 matched random assignments for each draw. Light and dark bars show the 1st–99th percentile range and interquartile range of the random scores. Diamonds show the fixed ChemOPD grouping.
Editing
Understanding
Method
Add P@1
Delete P@1
Sub. P@1
Murcko Sim.
Equiv. Acc.
FG MAE ↓
Ring MAE ↓
Ring-Sys Sim.
ChemOPD, Specialist-only
0.50
0.20
0.52
0.48
0.61
0.14
0.25
0.77
ChemOPD, no warm-up
0.45
0.85
0.60
0.50
0.57
0.09
0.20
0.73
ChemOPD
0.35
0.75
0.70
0.58
0.62
0.10
0.15
0.83
Appendix
Table 11: Component ablations for ChemOPD. Each panel follows the native metrics of Table 1 . FG and Ring are lower-is-better.
Figure 9: Capability trajectories and teacher-relative signals during OPD training. Panels (a–c) report the six task-averaged summaries defined in Equation 20 . Line style distinguishes the two summaries in each panel. Panel (d) reports 10-step means of the anchor and selected-specialist discrepancies in Equation 22 . The discrepancy curves are computed on each method’s own on-policy prefixes.
Figure 10: Representative test-set examples across the four chemical task families. Cases 1–6, marked with blue headers, show examples in which ChemOPD produces the better evaluated result than MOPD. Cases 7–10, marked with orange headers, show examples in which the comparison method performs better; Case 8 uses MOPD, while the other three use benchmark-family OPD. Orange highlights indicate reference-consistent structural features, and rose highlights indicate incorrect structural changes. The reported ESOL scores are calculated by the same property evaluator used for benchmark scoring.