cs.LGSep 27, 2026

ChemOPD: Multi-Teacher On-Policy Distillation for Multi-Task Chemical Reasoning

Authors: Yaoyao Xu, Xinjian Zhao, Xiaozhuang Song, Xuemin Chen, Tianshu Yu

Organizations: School of Data Science, The Chinese University of Hong Kong, Shenzhen · Shanghai Artificial Intelligence Laboratory

Abstract

Large language models are increasingly expected to support diverse chemical reasoning capabilities within a unified model. One approach is to develop specialized capabilities separately and consolidate them through multi-teacher on-policy distillation, but this raises two questions: how should specialization be organized, and how should specialist guidance be integrated? We introduce ChemOPD, which addresses both. We estimate task affinities from supervised fine-tuning gradients and solve a constrained mixed-integer program(MIP) to construct partially overlapping specialist groups. During distillation, we retain a generalist teacher trained on all tasks so that specialist guidance supplements rather than replaces its supervision. Our anchor-residual objective gradually increases the routed specialist's contribution on student-generated responses. On ChemCoTBench, affinity-guided specialization produces task-dependent gains over the generalist teacher and improves several capabilities beyond semantic task grouping. Yet stronger teacher-side performance does not automatically yield stronger students: with the same specialists and routes, anchor-residual OPD improves most reported metrics over specialist-only distillation and realizes a larger share of the available teacher gains. These results highlight specialization and capability integration as connected but distinct design problems in chemical reasoning.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Who Teaches Which Token? Verifier-Gated Multi-Expert On-Policy Distillation for Scientific Reasoning

    Sep 14, 2026Xun Xu, Zaixi Zhang

  2. TeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher Alignment

    Sep 27, 2026Zhenyu Lei, Zihan Chen, Yaochen Zhu +5Teacher-Student DistillationTeacher

  3. Structural Rationale Distillation via Reasoning Space Compression

    May 8, 2026Jialin Yang, Jiankun Wang, Jiajun Wu +3Reasoning PathsDataset Distillation