In many industrial applications, 1) tabular data is scarce and imbalanced and thus requires synthetic expansion; 2) input distributions drift between training and deployment (covariate shift); 3) validation sets often diverge from unseen test environments; or 4) standard generative models simply mimic outdated source distributions. This learning setting limits the stability of standard augmentation and adaptation pipelines. We generalize the task under such setting as the Augmented and Weighted Learning under Covariate Shift problem (AWL-CS). AWL-CS imposes two critical challenges on existing methods: 1) misleading generative guidance where models optimize for source similarity rather than downstream task relevance, and 2) structural instability of distributional density where reweighting mechanisms overfit to noisy validation signals. To tackle these challenges, we propose IGDPR (Invariant-Guided Diffusion with Prototype Reweighting), a unified framework that synergizes stable synthesis and structural adaptation: i) To achieve task-relevant generation, we steer the diffusion sampling process using invariant potentials to ensure synthetic samples align with stable decision boundaries rather than outdated correlations. ii) To ensure stable adaptation, we develop a prototype-based reweighting strategy that assesses sample reliability through structural clusters instead of isolated points, effectively filtering validation noise. Extensive experiments on real data demonstrate our method improves data quality by augmenting the most beneficial data for robust learning.
Figures & tables
Fig. 1 : Motivation of Augmented and Weighted Learning under Covariate Shift (AWL-CS). Under distribution shift, validation-guided augmentation and reweighting may amplify spurious patterns instead of improving generalization.
Fig. 2 : Overview of the proposed IGDPR (Invariant-Guided Diffusion with Prototype Reweighting) framework. The pipeline consists of four tightly coupled stages. (1) Invariant embedding construction: training data are partitioned into multiple pseudo-environments based on domain features, and a task-invariant latent space is learned via a multi-head TI-VAE with task, density, and invariance supervision. (2) Invariant-guided latent diffusion generation: a latent diffusion model is trained in the invariant space and, during sampling, its denoising trajectories are actively steered by frozen invariant guidance signals toward task-aligned and structurally stable regions. (3) Prototype-based reweighting: real and synthetic samples are clustered into structural prototypes, where reliability scores are estimated at the prototype level to avoid noisy and unstable point-wise reweighting. (4) Weighted downstream training: the downstream predictor is trained on the weighted augmented dataset to improve generalization under covariate shift.
Dataset
Source
Task Type
Samples
Feat Dim.
C-MAPSS
Simulated
Regression
∼ 53k
14
Scania APS
Real
Classification
76k
170
Gas Sensor
Real
Classification
13k
128
MetroPT-3
Real
Regression
1.5M
∼ 5
TABLE I : Summary of Datasets and Industrial Tasks.
Generator
Reweighter
Gas Sensor (Classif.)
NASA C-MAPSS (Reg.)
Scania ComponentX (Classif.)
MetroPT-3 (Reg.)
Acc ↑
F1-M ↑
F1-W ↑
MSE ↓
RMSE ↓
R2↑
Acc ↑
F1-M ↑
F1-W ↑
MSE ↓
RMSE ↓
R2↑
No Aug
Uniform
0.6479
0.6066
0.6069
2069.21
45.49
0.82
0.9083
0.8710
0.9021
0.2605
0.5104
0.71
Importance
0.6600
0.6225
0.6542
2067.63
45.47
0.83
0.9104
0.8748
0.9046
0.2589
0.5088
0.72
Approx.
0.4414
0.3921
0.4356
2093.73
45.76
0.78
0.8897
0.8415
0.8823
0.2516
0.5016
0.73
Prototype
0.6615
0.6284
0.6571
2064.41
45.44
0.83
0.9146
0.8793
0.9085
0.2598
0.5097
0.72
MLP GAN
Uniform
0.6410
0.6032
0.6354
2117.05
46.01
0.75
0.9121
0.8769
0.9068
0.2521
0.5021
0.73
TABLE II : Main Results on Industrial Domain Shift Benchmarks. We compare different generative augmentation methods combined with various reweighting strategies. We report Accuracy , F1-Macro (F1-M), and F1-Weighted (F1-W) for classification tasks, and MSE , RMSE , and R2 for regression tasks. The Guided Latent Diffusion generator denotes our invariant-guided latent diffusion module; paired with the Prototype reweighter (bold row) it constitutes our full method IGDPR. Bold indicates best performance.
Hyper-parameter
Scania
C-MAPSS
Gas Sensor
MetroPT-3
Batch Size
128
256
64
128
Learning Rate
1e-3
5e-4
1e-3
1e-3
Latent Dim ( z )
16
16
16
16
Prototypes ( K )
20
20
20
10
Reweight β
1.0
0.5
1.0
1.0
Reweight γ
1.0
0.5
0.5
1.0
TABLE III : Per-dataset hyperparameter configuration.
Fig. 3 : Ablation Study on Backbone, Reweighting Strategy, and Prototype Count. The bar charts (bottom axis) compare different generative backbones and reweighting strategies. The hatched bars denote our full method IGDPR (Guided Latent Diffusion + Prototype). The red lines (top axis) indicate the sensitivity to the number of prototypes K on a logarithmic scale. Metrics are Accuracy for the classification datasets (Gas Sensor, Scania), RMSE for C-MAPSS, and MSE for MetroPT-3.
Method
Gas Sensor
C-MAPSS
Scania
MetroPT-3
Acc
F1-M
F1-W
MSE
RMSE
R2
Acc
F1-M
F1-W
MSE
RMSE
R2
No Aug + Uniform
0.006
0.008
0.007
42.1
0.46
0.012
0.004
0.006
0.005
0.006
0.005
0.015
TabDDPM + Prototype
0.005
0.006
0.006
38.5
0.41
0.010
0.003
0.005
0.004
0.005
0.004
0.012
Guided Latent Diffusion + Prototype (IGDPR)
0.004
0.005
0.005
35.2
0.38
0.009
0.003
0.004
0.004
0.004
0.003
0.010
TABLE IV : Standard deviation across five random seeds , corresponding to the entries in Table II . Gains in the main table consistently exceed seed-level variance.
Fig. 4 : Sensitivity Analysis of Hyperparameters β and γ . We report Accuracy for Gas Sensor, F1-Macro for Scania, and MSE for C-MAPSS and MetroPT-3. Brighter colors (yellow) indicate better performance. For MSE, the color scale is inverted so that lower error corresponds to brighter colors. The optimal performance is consistently observed near β=1.0 and γ∈[0.5,1.0] .
Fig. 5 : t-SNE visualization of latent representations with generation and reweighting. Colors denote class labels, marker size indicates sample importance, and cross-shaped markers represent generated samples. Generated points that align with the original data manifold receive higher weights and help fill sparse regions, while off-manifold samples are assigned low importance.
Fig. 6 : Information Plane Analysis. Trade-off between Domain Information I(Z;E) and Task Information I(Z;Y) across datasets. “Ours” (marked IGD in the legend) denotes IGDPR, which consistently moves representations toward the ideal region (top-left), achieving lower domain dependence and higher task relevance compared to baselines.
λinv
0.0
0.1
0.5
1.0
2.0
10.0
RMSE
111.48
110.27
108.82
107.60
108.99
108.85
Δ vs λ=0
—
−1.21
−2.66
−3.88
−2.49
−2.63
TABLE V : λinv sweep on C-MAPSS (RMSE ↓ , single seed). The 5-seed validation at the optimum λinv=1.0 yields 108.74±0.96 RMSE, indicating that the sweep range ( ΔRMSE=3.88 ) is approximately 4× the seed-level noise.
λinv
0.0
0.1
0.5
1.0
2.0
10.0
RMSE
0.583
0.509
0.486
0.578
0.476
0.543
Δ vs λ=0
—
−0.074
−0.097
−0.005
−0.107
−0.040
TABLE VI : λinv sweep on MetroPT-3 (RMSE ↓ , single seed). The 5-seed validation at λinv=0.1 yields 0.510±0.029 RMSE; seed noise is comparable to inter- λ gaps, so single-seed rankings within the [0.1,2.0] plateau should be interpreted as a low-error region rather than a strict ordering.
Dataset
Env. construction
I(E;X)max
I(E;X)Σ
I(E;Y)
I(E;Y)−I(E;Y^)
Scania
residual K -means ( K=3 )
0.396
4.757
0.000
−0.002
C-MAPSS
explicit op. condition
0.501
0.954
0.007
−0.000
Gas Sensor
explicit batch id
0.591
7.790
0.132
−0.003
MetroPT-3
K -means ( K=3 ) on PCA(5) features
0.626
1.146
0.469
−0.161
TABLE VII : MI probe of per-dataset environment constructions. All four clear the I(E;X)max≥0.05 nats identifiability floor by at least an order of magnitude; the non-positive residual I(E;Y)−I(E;Y^) is consistent with env–label dependence being mediated by X .