Supervised fine-tuning performance for large language models depends strongly on how training budget is distributed across a heterogeneous set of tasks. In practice, mixtures are often fixed using simple heuristics (e.g., uniform or size-proportional sampling) that ignore task interactions, which can hurt transfer and waste budget on redundant sources. We introduce TaskPGM, a framework for learning continuous task mixtures via an energy-based model over tasks. Tasks form the nodes of a Markov random field: unary potentials capture per-task utility, and pairwise potentials encode inter-task relationships using behavioral divergences computed from predictive distributions of single-task fine-tuned models (e.g., Jensen--Shannon divergence and pointwise mutual information). Optimizing this objective yields mixtures that balance coverage against redundancy. We show that the resulting set function is weakly submodular under budget constraints, enabling approximation guarantees for discrete selection variants. Across multiple model families (LLaMA-7B, Qwen2-7B) and evaluation suites (BIG-Bench Hard), TaskPGM improves over standard mixing strategies and provides interpretable structure over task interactions.
Figure 2 : MRF over tasks with similarities Sij (left) and the learned mixture probabilities pi∗ from Eq. ( 2 ) (right; shown the corr. probabilities).
Figure 3 : Accuracy surfaces over β/λ ratios during greedy mixture construction (LLama-7B) For each downstream benchmark, the surface plots evaluation accuracy as a function of the greedy step k (mixture size) and the weight ratio β/λ , with β and λ varied on a log scale, illustrating how performance curvature varies across tasks.
Size
Method
Leaderboard
BBH
GPQA
IFEval
Math
MMLU-Pro
MUSR
25K
Random
0.4909±0.0062
0.3188±0.1350
0.3285±0.0081
0.1881±0.0100
0.4157±0.0045
0.4603±0.0180
25K
Uniform
0.5013±0.0062
0.3314±0.0136
0.2926±0.0074
0.2085±0.0104
0.4161±0.0045
0.4683±0.0180
25K
EPM
0.4970±0.0062
0.3314±0.0136
0.3094±0.0077
0.2068±0.0105
0.4146±0.0045
0.4537±0.0180
25K
LESS
0.5173±0.0062
0.3020±0.0094
0.3297±0.0080
0.2002±0.0101
0.4230±0.0045
0.3598±0.0170
25K
Ours (PMI)
0.5202 ±0.0062
0.3341 ±0.0136
0.3909±0.0085
0.2096±0.0101
0.4251 ±0.0045
0.4701 ±0.0180
Table 1 : Qwen2-7B: Instruct-tuning performance on Leaderboard (25K and 50K samples).
Size
Method
Leaderboard
BBH
GPQA
IFEval
Math
MMLU-Pro
MUSR
25K
Random
0.3482 ±0.0059
0.2626 ±0.0128
0.3465 ±0.0000
0.0098 ±0.0027
0.1877 ±0.0036
0.3677 ±0.0172
25K
Uniform
0.3501 ±0.0059
0.2701 ±0.0129
0.3501 ±0.0000
0.0151 ±0.0034
0.1768 ±0.0035
0.4027 ±0.0175
25K
EPM
0.3593 ±0.0059
0.2601 ±0.0127
0.3405 ±0.0000
0.0151 ±0.0033
0.1836 ±0.0035
0.4286 ±0.0177
25K
LESS
0.4059±0.0055
0.2412±0.0088
0.3417±0.0000
0.0001±0.0004
0.1816±0.0062
0.4157±0.0175
25K
Ours (PMI)
0.4095 ±0.0054
0.2718 ±0.0129
0.3561 ±0.0000
0.0159 ±0.0034
0.1924 ±0.0036
0.4298 ±0.0177
Table 2 : Llama2-7B: Instruct-tuning performance on Leaderboard (25K and 50K samples).
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4 : Comparison of task similarity metrics using cosine similarity (top), and PMI and JSD-based heatmaps (bottom). Cosine scores are generally low and fail to distinguish task structure. PMI highlights asymmetric task alignment. JSD offers symmetric, bounded divergence and reveals clearer task groupings across models.
Figure 5 : Task similarity heatmaps for Qwen models computed using Jensen–Shannon divergence (left) and pointwise mutual information (right).
Figure 6 : Eigenvalue spectra of similarity matrices derived from (a) Jensen–Shannon divergence (JSD) and (b) pointwise mutual information (PMI) . The PMI-based matrix exhibits a steeper spectral decay , indicating a lower effective rank and thus a more compact embedding of similarity relationships.
Dataset
Leaderboard
Size / Method
BBH
GPQA
IFEval
Math
MMLU-Pro
MUSR
25K
Random
0.3482 ±0.0059
0.2626 ±0.0128
0.3465 ±N/A
0.0098 ±0.0027
0.1877 ±0.0036
0.3677 ±0.0172
Uniform
0.3501 ±0.0059
0.2701 ±0.0129
0.3501 ±N/A
0.0151 ±0.0034
0.1768 ±0.0035
0.4027 ±0.0175
EPM
0.3593 ±0.0059
0.2601 ±0.0127
0.3405 ±N/A
0.0151 ±0.0033
0.1836 ±0.0035
0.4286 ±0.0177
LESS
0.4059±0.0055
0.2412±0.0088
0.3417±N/A
0.0001±0.0004
0.1816±0.0062
0.4157±0.0175
Appendix
Table 3 : Llama-2-7b: Instruction-tuning performance on Leaderboard subsets with β=20,λ=10 using batch size 8.
Dataset Size(Method)
Leaderboard
BBH
GPQA
IFEval
Math
MMLU-Pro
MUSR
25K (PMI)
β =14954 ; λ =263
0.3637 ±0.0059
0.2685 ±0.0128
0.3405 ±N/A
0.0159 ±0.0034
0.1869 ±0.0036
0.3849 ±0.0173
β =5273 ; λ =195
0.3536 ±0.0059
0.2718 ±0.0129
0.3609 ±N/A
0.0166 ±0.0035
0.1823 ±0.0035
0.4021 ±0.0174
β =2535 ; λ =196
0.3659 ±0.0059
0.2701 ±0.0129
0.3357 ±N/A
0.0128 ±0.0031
0.1890 ±0.0036
0.3929 ±0.0173
β =307 ; λ =60
0.3605 ±0.0059
0.2735 ±0.0129
0.3681 ±N/A
0.0159 ±0.0034
0.1881 ±0.0036
0.4074 ±0.0174
Appendix
Table 4 : Llama-2-7b: Instruction-tuning performance on Leaderboard subsets with varying β and λ using batch size 8.
Fine-tuning Multimodal Large Language Models (MLLMs) on specialized tasks often leads to catastrophic forgetting of their general capabilities. Existing model merging methods to combat this are often heuristic or use sub-optimal objectives. We propose CurvatureGuided Mixing (CGM), a theoretically grounded framework that merges pre-trained and fine-tuned models. CGM formulates a joint optimization objective and uses a second-order (Hessian) approximation of the loss landscapes to analytically derive an optimal, closed-form "soft mixing" ratio. This ratio intelligently blends parameters based on their relative task-specific curvatures. We also introduce CGM†, a robust "hard mixing" variant that performs sparse parameter selection guided by a novel, curvature-aware score. Experiments on LLaVA-1.5 and Qwen2.5VL across multiple downstream tasks show that CGM and CGM† consistently improve the trade-off between task specialization and general knowledge retention over existing methods. Code is available at github.com/zzsyjl/CGM-ECCV-2026.
Jinglong Yang, Jiaxuan He, Wenjian Huang +2
Research Institute of Trustworthy Autonomous Systems and Department of Computer Science and Engineering, Southern University of Science and Technology · Department of Computer Science, City University of Hong Kong
Adapting a language model to a specialized corpus means choosing which instruction-tuning tasks to train on under a fixed budget, and testing one choice costs a fine-tuning run. Common heuristics add more source tasks or pick sources similar to the target. The first assumes transfer is never negative; the second, that it is symmetric. We show that both assumptions fail: task A can help task B while B hurts A, so helpfulness is a signed property of ordered source--target pairs. We introduce the transfer map, a signed estimate of how much each source helps or hurts each held-out target. We fit the map in hundreds of fine-tuning runs on Qwen3 and Mistral models from 0.6B to 32B parameters, with all sources drawn from one corpus and no training examples from the target. The map predicts a held-out target's accuracy on unseen mixtures: recorded before those runs, its predictions have less than half the error of a mixture-agnostic baseline. The map is specific to its target and corpus but transfers across model scale: a mixture selected in advance at one size beats training on all source tasks at every other size we tested. Transfer is thus a property of the data. The map selects the tasks that help and drops the one that interferes: accuracy on the reasoning targets (causal explanation, multi-hop questions and methodological critique) rises by up to 14 percentage points over training on all source tasks.
Nima H. Siboni, Vahid Rostami
Juna.ai, Kastanienallee 32, 10435 Berlin, Germany · Computational Systems Neuroscience, Institute of Zoology, University of Cologne, Germany
Not all training samples contribute equally to large language model fine-tuning. Selecting informative training samples can reduce the computational cost while preserving downstream performance. Many existing data selection methods rely on indirect heuristics, such as data quality, diversity or reasoning trace length. However, the effectiveness of these fixed criteria is task-dependent and difficult to generalize across diverse downstream tasks. Perplexity-based data selection provides a simple and model-aware solution to estimate the sample difficulty, but existing approaches typically score the entire training sequence and ignore the difference in learning objectives of language modeling and reasoning tasks. In this paper, we propose PPL-Factory, a simple and interpretable data selection framework that combines task-aware perplexity-based scores and data budget-aware selection criteria. Experiments on GSM8K demonstrate that PPL-Factory outperforms other state-of-the-art data selection methods using only 1% of the training set. With 10% of the data, PPL-Factory exceeds full-data fine-tuning accuracy by 0.9 on GSM8K and 4.8 on MATH. Overall, our results demonstrate that task-aware and budget-aware perplexity-based selection provides an effective and applicable approach for efficient fine-tuning.
Hang Zhang, Warren J. Gross
Department of Electrical and Computer Engineering, McGill University Montreal, QC, Canada