Meta-learning aims to leverage information across related tasks to improve prediction on unlabeled data for new tasks when only a small number of labeled observations are available ("few-shot" learning). Increased task diversity is often believed to enhance meta-learning by providing richer information across tasks. However, recent work by Kumar et al. (2022) shows that increasing task diversity, quantified through the overall geometric spread of task representations, can in fact degrade meta-learning prediction performance across a range of models and datasets. In this work, we build on this observation by showing that meta-learning performance is affected not only by the overall geometric variability of task parameters, but also by how this variability is allocated relative to an underlying low-dimensional structure. Similar to Pimonova et al. (2025), we decompose task-specific regression effects into a structurally informative component and an orthogonal, non-informative component. We show theoretically and through simulation that meta-learning prediction degrades when a larger fraction of between-task variability lies in orthogonal, non-informative directions, even when the overall geometric variability of tasks is held fixed.
Figures & tables
Figure 1 : This figure displays the density of log(sin2(θ1)) , representing the distance between the true P0 and posterior samples of P for different values of φ0 .
Figure 2 : This figure on the top presents the density of R2 values across 100 datasets with n=50 data points, comparing meta-learning prediction for tasks generated with φ0∈{0.2,0.15,0.1,0.05,0.02,0.01} . The figure in the bottom presents the density of trace(Σy) values across 100 datasets, comparing uncertainty in meta-learning prediction for tasks generated from various φ0 .
φ0
R2
trace(Σy)
0.20
0.6492
242.0127
0.15
0.6886
193.3547
0.10
0.7258
137.8519
0.05
0.8410
84.1290
0.02
0.8736
39.2434
0.01
0.9157
25.1929
Table 1 : Aggregate simulation results across different values of φ0 .
Figure 3 : This figure displays the density of log(sin2(θ1)) , representing the distance between the true P0 and posterior samples of P for different pairs of (φ0,k) with k/trace(Σ0)=0.169(dotted),0.423(dashed),0.847(solid) , where trace(Σ0)=11.8 ,
Figure 4 : This figure on the top presents the density of R2 values across 100 datasets with n=50 data points, comparing meta-learning prediction for tasks generated using (φ0,k)=(0.1,2),(0.05,5),(0.02,10) with corresponding k/trace(Σ0)=0.169(dotted),0.423(dashed),0.847(solid) . The figure in the bottom presents the density of trace(Σy) values across the same datasets, under the same task generation settings.
Figure 5 : Logarithm of sin2(θ1) are plotted on the x -axis and the density of the values are plotted on the y -axis. This figure illustrates the decline of sin2θ1(P[t],P⋆) as the number of tasks S and the number of samples per task ns increase, under a high-dimensional setting with ns=50 (red) and a moderate-dimensional setting with ns=100 (black) samples per task.
Figure 6 : This plot presents the density of R2 values from meta-learning models based on the posterior distribution of the meta-parameters P and φ , estimated from meta-training with 100 (solid), 500 (dashed), and 2000 (dotted) tasks, each task containing either 50 (red) or 100 (black) samples. In the meta-test phase, β⋆ is updated using 70 training samples from a new task, and predictions are evaluated on 30 additional samples from the same task using both meta-learning models and LASSO(blue).
Figure 7 : This figure displays the variance of the posterior predictive distribution of y , obtained by training β⋆ using 70 training samples in the meta-testing stage and evaluated on 30 validation samples.
Meta-learning represents a strong class of approaches for solving few-shot learning tasks. Nonetheless, recent research suggests that simply pre-training a generic encoder can potentially surpass meta-learning algorithms. In this paper, we first discuss the reasons why meta-learning fails to stand out in these few-shot learning experiments, and hypothesize that it is due to the few-shot learning tasks lacking diversity. We propose DRESS, a task-agnostic Disentangled REpresentation-based Self-Supervised meta-learning approach that enables fast model adaptation on highly diversified few-shot learning tasks. Specifically, DRESS utilizes disentangled representation learning to create self-supervised tasks that can fuel the meta-training process. Furthermore, we also propose a class-partition based metric for quantifying the task diversity directly on the input space. We validate the effectiveness of DRESS through experiments on datasets with multiple factors of variation and varying complexity. The results suggest that DRESS is able to outperform competing methods on the majority of the datasets and task setups. Through this paper, we advocate for a re-examination of proper setups for task adaptation studies, and aim to reignite interest in the potential of meta-learning for solving few-shot learning tasks via disentangled representations.
Accurate bioactivity prediction is a central challenge in early-stage drug discovery, as individual assays often contain too few measurements to train reliable models independently. Meta-learning offers a principled approach to this few-shot setting, but assay heterogeneity may limit its effectiveness. Here, we test this hypothesis and show that meta-learning performance degrades as meta-training tasks become more heterogeneous. To address this, we introduce MetaHeta, a meta-learning framework that accounts for assay heterogeneity by conditioning predictions on auxiliary data from related assays, with relatedness defined flexibly from available assay information. The architecture of MetaHeta combines linear attention over large auxiliary datasets with exact attention over scarce task-specific context, enabling efficient scaling to the former without compromising exact attention over the latter. We demonstrate the benefits of our approach on assays from ChEMBL and BindingDB, improving few-shot bioactivity prediction and downstream compound prioritization in retrospective Bayesian optimization.
Michal Kmicikiewicz, Tommy Rochussen, Vincent Fortuin +1
Institute of AI for Health, Helmholtz Munich · School of Computation, Information and Technology, Technical University of Munich · Munich Center for Machine Learning +4
In multi-task learning (MTL) negative transfer is often considered as an optimization artifact, but it can also be viewed as a consequence of limited shared capacity and weak task redundancy. We investigate this effect through a Capacity--Redundancy (CR) identity that decomposes the sum of per-task predictive informations into joint predictive information that includes label redundancy defined via total correlation (TC), and a residual coupling term that quantifies interference left unresolved by the shared representation. Additionally, we show two key results: (i) a clustering-gap decomposition that gives a necessary and sufficient condition for clustered sharing to outperform global sharing, and (ii) a gradient--TC bridge in a Gaussian multi-task model that formally justifies gradient cosine similarity as a proxy for redundancy ordering. Empirically, we estimate the residual coupling Δ from validation residual correlations, showing that clustered LoRA substantially reduces Δ, outperforms size-matched random partitions, and results in statistically significant gains with multi-seed confidence intervals.