Supervised fine-tuning performance for large language models depends strongly on how training budget is distributed across a heterogeneous set of tasks. In practice, mixtures are often fixed using simple heuristics (e.g., uniform or size-proportional sampling) that ignore task interactions, which can hurt transfer and waste budget on redundant sources. We introduce TaskPGM, a framework for learning continuous task mixtures via an energy-based model over tasks. Tasks form the nodes of a Markov random field: unary potentials capture per-task utility, and pairwise potentials encode inter-task relationships using behavioral divergences computed from predictive distributions of single-task fine-tuned models (e.g., Jensen--Shannon divergence and pointwise mutual information). Optimizing this objective yields mixtures that balance coverage against redundancy. We show that the resulting set function is weakly submodular under budget constraints, enabling approximation guarantees for discrete selection variants. Across multiple model families (LLaMA-7B, Qwen2-7B) and evaluation suites (BIG-Bench Hard), TaskPGM improves over standard mixing strategies and provides interpretable structure over task interactions.
Figure 2 : MRF over tasks with similarities Sij (left) and the learned mixture probabilities pi∗ from Eq. ( 2 ) (right; shown the corr. probabilities).
Figure 3 : Accuracy surfaces over β/λ ratios during greedy mixture construction (LLama-7B) For each downstream benchmark, the surface plots evaluation accuracy as a function of the greedy step k (mixture size) and the weight ratio β/λ , with β and λ varied on a log scale, illustrating how performance curvature varies across tasks.
Size
Method
Leaderboard
BBH
GPQA
IFEval
Math
MMLU-Pro
MUSR
25K
Random
0.4909±0.0062
0.3188±0.1350
0.3285±0.0081
0.1881±0.0100
0.4157±0.0045
0.4603±0.0180
25K
Uniform
0.5013±0.0062
0.3314±0.0136
0.2926±0.0074
0.2085±0.0104
0.4161±0.0045
0.4683±0.0180
25K
EPM
0.4970±0.0062
0.3314±0.0136
0.3094±0.0077
0.2068±0.0105
0.4146±0.0045
0.4537±0.0180
25K
LESS
0.5173±0.0062
0.3020±0.0094
0.3297±0.0080
0.2002±0.0101
0.4230±0.0045
0.3598±0.0170
25K
Ours (PMI)
0.5202 ±0.0062
0.3341 ±0.0136
0.3909±0.0085
0.2096±0.0101
0.4251 ±0.0045
0.4701 ±0.0180
Table 1 : Qwen2-7B: Instruct-tuning performance on Leaderboard (25K and 50K samples).
Size
Method
Leaderboard
BBH
GPQA
IFEval
Math
MMLU-Pro
MUSR
25K
Random
0.3482 ±0.0059
0.2626 ±0.0128
0.3465 ±0.0000
0.0098 ±0.0027
0.1877 ±0.0036
0.3677 ±0.0172
25K
Uniform
0.3501 ±0.0059
0.2701 ±0.0129
0.3501 ±0.0000
0.0151 ±0.0034
0.1768 ±0.0035
0.4027 ±0.0175
25K
EPM
0.3593 ±0.0059
0.2601 ±0.0127
0.3405 ±0.0000
0.0151 ±0.0033
0.1836 ±0.0035
0.4286 ±0.0177
25K
LESS
0.4059±0.0055
0.2412±0.0088
0.3417±0.0000
0.0001±0.0004
0.1816±0.0062
0.4157±0.0175
25K
Ours (PMI)
0.4095 ±0.0054
0.2718 ±0.0129
0.3561 ±0.0000
0.0159 ±0.0034
0.1924 ±0.0036
0.4298 ±0.0177
Table 2 : Llama2-7B: Instruct-tuning performance on Leaderboard (25K and 50K samples).
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4 : Comparison of task similarity metrics using cosine similarity (top), and PMI and JSD-based heatmaps (bottom). Cosine scores are generally low and fail to distinguish task structure. PMI highlights asymmetric task alignment. JSD offers symmetric, bounded divergence and reveals clearer task groupings across models.
Figure 5 : Task similarity heatmaps for Qwen models computed using Jensen–Shannon divergence (left) and pointwise mutual information (right).
Figure 6 : Eigenvalue spectra of similarity matrices derived from (a) Jensen–Shannon divergence (JSD) and (b) pointwise mutual information (PMI) . The PMI-based matrix exhibits a steeper spectral decay , indicating a lower effective rank and thus a more compact embedding of similarity relationships.
Dataset
Leaderboard
Size / Method
BBH
GPQA
IFEval
Math
MMLU-Pro
MUSR
25K
Random
0.3482 ±0.0059
0.2626 ±0.0128
0.3465 ±N/A
0.0098 ±0.0027
0.1877 ±0.0036
0.3677 ±0.0172
Uniform
0.3501 ±0.0059
0.2701 ±0.0129
0.3501 ±N/A
0.0151 ±0.0034
0.1768 ±0.0035
0.4027 ±0.0175
EPM
0.3593 ±0.0059
0.2601 ±0.0127
0.3405 ±N/A
0.0151 ±0.0033
0.1836 ±0.0035
0.4286 ±0.0177
LESS
0.4059±0.0055
0.2412±0.0088
0.3417±N/A
0.0001±0.0004
0.1816±0.0062
0.4157±0.0175
Appendix
Table 3 : Llama-2-7b: Instruction-tuning performance on Leaderboard subsets with β=20,λ=10 using batch size 8.
Dataset Size(Method)
Leaderboard
BBH
GPQA
IFEval
Math
MMLU-Pro
MUSR
25K (PMI)
β =14954 ; λ =263
0.3637 ±0.0059
0.2685 ±0.0128
0.3405 ±N/A
0.0159 ±0.0034
0.1869 ±0.0036
0.3849 ±0.0173
β =5273 ; λ =195
0.3536 ±0.0059
0.2718 ±0.0129
0.3609 ±N/A
0.0166 ±0.0035
0.1823 ±0.0035
0.4021 ±0.0174
β =2535 ; λ =196
0.3659 ±0.0059
0.2701 ±0.0129
0.3357 ±N/A
0.0128 ±0.0031
0.1890 ±0.0036
0.3929 ±0.0173
β =307 ; λ =60
0.3605 ±0.0059
0.2735 ±0.0129
0.3681 ±N/A
0.0159 ±0.0034
0.1881 ±0.0036
0.4074 ±0.0174
Appendix
Table 4 : Llama-2-7b: Instruction-tuning performance on Leaderboard subsets with varying β and λ using batch size 8.
Research Institute of Trustworthy Autonomous Systems and Department of Computer Science and Engineering, Southern University of Science and Technology · Department of Computer Science, City University of Hong Kong