Robust LassoNet: Enhancing Feature Selection in Neural Networks via Robust Loss Functions
Authors: Daniela De Canditiis, Italia De Feis, Paola Stolfi
Organizations: Istituto per le Applicazioni del Calcolo “Mauro Picone”, Consiglio Nazionale delle Ricerche, Via dei Taurini, 19, Roma, 00185, Italia · Istituto per le Applicazioni del Calcolo “Mauro Picone”, Consiglio Nazionale delle Ricerche, Via Pietro Castellino, 111, Napoli, 80131, Italia
Feature selection in neural networks remains a challenging problem, particularly in the presence of noisy or contaminated data. LassoNet is a recent approach that addresses this issue by combining neural networks with hierarchical sparsity constraints, enabling simultaneous prediction and variable selection. However, its standard formulation relies on the mean squared error (MSE) loss, which is known to be highly sensitive to outliers. In this paper, we present Robust LassoNet, an extension of LassoNet that incorporates robust loss functions, such as Huber, Cauchy, Tukey's bisquare, and Nonnegative Garrote, to mitigate the effect of extreme observations. The proposed approach preserves the original optimization framework while improving stability under data contamination. Through experiments on synthetic and real datasets, we show that robust LassoNet significantly improves both predictive performance and feature selection accuracy in the presence of outliers or heavy tailed noise, while maintaining comparable performance in clean settings. These results highlight the importance of robustness in neural network based feature selection and suggest practical guidelines for choosing appropriate loss functions.
Figures & tables
Parameter
Default
Description
hidden_dims
(100,)
Dimensions of the hidden layers
lambda_start
’auto’
Initial value of λ along the path
lambda_seq
None
Explicit sequence of λ values
γ
0.0
L2 regularization applied to network weights
γskip
0.0
L2 regularization applied to the skip connection
path_multiplier
1.02
Multiplicative factor along the path ( λ←λ⋅c )
Table 1: Parameters of LassoNetRegressor .
Gaussian contamination
Moderate mean contamination
Strong mean contamination
Loss
PE
NSR
FPR
FNR
PE
NSR
FPR
FNR
PE
NSR
FPR
FNR
mse
0.398
0.122
0.103
0.152
0.532
0.116
0.098
0.184
0.850
0.046
0.030
0.320
huber
0.399
0.087
0.068
0.176
0.348
0.062
0.043
0.192
0.347
0.034
0.015
0.200
cauchy
0.402
0.102
0.083
0.160
0.342
0.072
0.053
0.192
0.331
0.061
0.042
0.192
tukey
0.400
0.091
0.073
0.176
0.320
0.066
0.047
0.192
0.308
0.071
0.052
0.200
garrote
0.400
0.090
0.071
0.176
0.338
0.058
0.039
0.192
0.331
0.042
0.023
0.200
Table 2: Performance under Gaussian and mean contamination scenarios for Model 1 .
Gaussian contamination
Moderate mean contamination
Strong mean contamination
Loss
PE
NSR
FPR
FNR
PE
NSR
FPR
FNR
PE
NSR
FPR
FNR
mse
0.253
0.180
0.155
0.000
0.282
0.131
0.105
0.007
0.320
0.055
0.028
0.067
huber
0.253
0.104
0.076
0.000
0.262
0.115
0.087
0.000
0.265
0.092
0.065
0.007
cauchy
0.264
0.129
0.102
0.000
0.266
0.139
0.112
0.000
0.268
0.111
0.084
0.007
tukey
0.255
0.106
0.078
0.000
0.263
0.130
0.103
0.000
0.265
0.103
0.076
0.007
garrote
0.257
0.122
0.095
0.000
0.274
0.140
0.114
0.007
0.286
0.100
0.072
0.007
Table 3: Performance under Gaussian and mean contamination scenarios for Model 2 .
Gaussian contamination
Moderate mean contamination
Strong mean contamination
Loss
PE
NSR
FPR
FNR
PE
NSR
FPR
FNR
PE
NSR
FPR
FNR
mse
0.498
0.266
0.244
0.027
0.539
0.145
0.125
0.227
0.894
0.077
0.060
0.367
huber
0.523
0.137
0.114
0.100
0.433
0.113
0.092
0.180
0.447
0.122
0.100
0.187
cauchy
0.532
0.106
0.082
0.107
0.428
0.108
0.086
0.180
0.418
0.118
0.096
0.160
tukey
0.536
0.142
0.118
0.094
0.435
0.105
0.083
0.193
0.397
0.090
0.066
0.154
garrote
0.510
0.152
0.129
0.080
0.444
0.103
0.081
0.200
0.403
0.124
0.101
0.154
Table 4: Performance under Gaussian and mean contamination scenarios for Model 3 .
Gaussian contamination
Moderate mean contamination
Strong mean contamination
Loss
PE
NSR
FPR
FNR
PE
NSR
FPR
FNR
PE
NSR
FPR
FNR
mse
0.282
0.188
0.146
0.000
1.841
0.018
0.009
0.816
3.300
0.005
0.004
0.992
huber
0.294
0.136
0.091
0.000
0.279
0.108
0.062
0.028
0.278
0.101
0.055
0.036
cauchy
0.295
0.154
0.109
0.000
0.238
0.131
0.086
0.012
0.277
0.097
0.055
0.108
tukey
0.296
0.153
0.109
0.004
0.240
0.136
0.091
0.000
0.240
0.136
0.091
0.000
garrote
0.291
0.141
0.096
0.000
0.287
0.097
0.057
0.148
0.640
0.108
0.067
0.124
Table 5: Performance under Gaussian and mean contamination scenarios for Model 4 .
Gaussian contamination
Moderate variance contamination
Strong variance contamination
Loss
PE
NSR
FPR
FNR
PE
NSR
FPR
FNR
PE
NSR
FPR
FNR
mse
0.398
0.122
0.103
0.152
0.407
0.135
0.118
0.192
0.626
0.036
0.025
0.536
huber
0.399
0.087
0.068
0.176
0.339
0.081
0.063
0.192
0.336
0.054
0.035
0.200
cauchy
0.402
0.102
0.083
0.160
0.339
0.081
0.062
0.192
0.330
0.072
0.053
0.184
tukey
0.400
0.091
0.073
0.176
0.336
0.091
0.072
0.184
0.323
0.078
0.059
0.192
garrote
0.400
0.090
0.071
0.176
0.338
0.079
0.060
0.192
0.326
0.081
0.062
0.184
Table 6: Performance under Gaussian and variance contamination scenarios for Model 1 .
Gaussian contamination
Moderate variance contamination
Strong variance contamination
Loss
PE
NSR
FPR
FNR
PE
NSR
FPR
FNR
PE
NSR
FPR
FNR
mse
0.253
0.180
0.155
0.000
0.321
0.147
0.123
0.067
0.639
0.034
0.021
0.573
huber
0.253
0.104
0.076
0.000
0.235
0.117
0.090
0.000
0.246
0.070
0.042
0.007
cauchy
0.264
0.129
0.102
0.000
0.237
0.130
0.103
0.000
0.229
0.134
0.107
0.000
tukey
0.255
0.106
0.078
0.000
0.228
0.114
0.086
0.000
0.220
0.084
0.056
0.000
garrote
0.257
0.122
0.095
0.000
0.236
0.123
0.096
0.000
0.227
0.133
0.107
0.000
Table 7: Performance under Gaussian and variance contamination scenarios for Model 2 .
Gaussian contamination
Moderate variance contamination
Strong variance contamination
Loss
PE
NSR
FPR
FNR
PE
NSR
FPR
FNR
PE
NSR
FPR
FNR
mse
0.498
0.266
0.244
0.027
0.687
0.126
0.107
0.280
1.242
0.019
0.015
0.833
huber
0.523
0.137
0.114
0.100
0.449
0.157
0.135
0.140
0.456
0.129
0.107
0.180
cauchy
0.532
0.106
0.082
0.107
0.444
0.133
0.111
0.160
0.436
0.141
0.119
0.160
tukey
0.536
0.142
0.118
0.094
0.450
0.126
0.104
0.154
0.433
0.113
0.090
0.147
garrote
0.510
0.152
0.129
0.080
0.432
0.154
0.132
0.134
0.417
0.122
0.099
0.134
Table 8: Performance under Gaussian and variance contamination scenarios for Model 3 .
Gaussian contamination
Moderate variance contamination
Strong variance contamination
Loss
PE
NSR
FPR
FNR
PE
NSR
FPR
FNR
PE
NSR
FPR
FNR
mse
0.282
0.188
0.146
0.000
0.350
0.173
0.136
0.120
0.666
0.032
0.021
0.756
huber
0.294
0.136
0.091
0.000
0.259
0.158
0.114
0.004
0.259
0.094
0.049
0.044
cauchy
0.295
0.154
0.109
0.000
0.259
0.160
0.115
0.000
0.251
0.147
0.103
0.004
tukey
0.296
0.153
0.109
0.004
0.260
0.171
0.128
0.000
0.249
0.145
0.099
0.000
garrote
0.291
0.141
0.096
0.000
0.258
0.154
0.109
0.004
0.247
0.137
0.092
0.000
Table 9: Performance under Gaussian and variance contamination scenarios for Model 4 .
Variable
Description
Possible outcomes
MPG
Fuel consumption (in miles per gallon)
Fuel
Type of fuel
Diesel, Petrol
Price
List price (in UK pounds)
Cylinders
Number of cylinders in the engine
Displacement
Displacement of the engine (in cc)
DriveWheel
Type of drive wheel
4WD, Front, Rear
Table 10: Description of the Top Gear car data.
Loss
trimmed PE
NSR
mse
0.570
0.734
huber
0.193
0.455
cauchy
0.197
0.500
tukey
0.204
0.508
garrote
0.186
0.511
rf
0.214
0.259
Table 11: Performance for Top Gear dataset. All metrics have been averaged over 50 independent runs.
Figure 1: Boxplots of performance metrics for the Top Gear dataset. The left panel reports trimmed PE, while the right panel reports NSR at the variable level, across the five methods.
Figure 2: Heatmap of variable selection for the Top Gear dataset. Rows correspond to the considered methods, and columns to variables selected in at least 70% of the 50 runs by at least one method. White cells indicate absence of selection, while colored cells denote selected variables. Color intensity reflects the level of agreement across methods, ranging from light yellow (selection by a single method) to red (selection by all eight methods).
Loss
trimmed PE
NSR
mse
10.356
0.278
huber
6.759
0.402
cauchy
6.598
0.474
tukey
6.535
0.631
garrote
6.532
0.528
rf
12.398
0.278
Table 12: Performance for Atherosclerosis dataset. All metrics have been averaged over 50 independent runs.
Figure 3: Boxplots of performance metrics for the Atherosclerosis dataset. The left panel reports trimmed PE, while the right panel reports NSR at the variable level, across the five methods.
Figure 4: Heatmap of variable selection for the Atherosclerosis dataset. Rows correspond to the considered methods, and columns to variables selected in at least 70% of the 50 runs by at least one method. White cells indicate absence of selection, while colored cells denote selected variables. Color intensity reflects the level of agreement across methods, ranging from light yellow (selection by a single method) to red (selection by all seven methods). Note that LassoNet with the MSE loss is not included, as no variable was selected in at least 70% of the 50 runs.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Boxplots of the performance metrics results for model 1 with 10% contamination. First column refers to gaussian scenario; second column to moderate mean scenario; third column to strong mean scenario. The first row shows a realization of the true and contaminated training response y for all the scenarios.
Figure 6: Boxplots of the performance metrics results for model 2 with 10% contamination. First column refers to gaussian scenario; second column to moderate mean scenario; third column to strong mean scenario. The first row shows a realization of the true and contaminated training response y for all the scenarios.
Figure 7: Boxplots of the performance metrics results for model 3 with 10% contamination. First column refers to gaussian scenario; second column to moderate mean scenario; third column to strong mean scenario. The first row shows a realization of the true and contaminated training response y for all the scenarios.
Figure 8: Boxplots of the performance metrics results for model 4 with 10% contamination. First column refers to gaussian scenario; second column to moderate mean scenario; third column to strong mean scenario. The first row shows a realization of the true and contaminated training response y for all the scenarios.
Figure 9: Boxplots of the performance metrics results for model 1 with 10% contamination. First column refers to gaussian scenario; second column to moderate variance scenario; third column to strong variance scenario. The first row shows a realization of the true and contaminated training response y for all the scenarios.
Figure 10: Boxplots of the performance metrics results for model 2 with 10% contamination. First column refers to gaussian scenario; second column to moderate variance scenario; third column to strong variance scenario. The first row shows a realization of the true and contaminated training response y for all the scenarios.
Figure 11: Boxplots of the performance metrics results for model 3 with 10% contamination. First column refers to gaussian scenario; second column to moderate variance scenario; third column to strong variance scenario. The first row shows a realization of the true and contaminated training response y for all the scenarios.
Figure 12: Boxplots of the performance metrics results for model 4 with 10% contamination. First column refers to gaussian scenario; second column to moderate variance scenario; third column to strong variance scenario. The first row shows a realization of the true and contaminated training response y for all the scenarios.
Learner
Parameter
Value
RandomForest
criterion
absolute_error
n_estimators
300
(all others)
scikit-learn defaults
XGBoost
objective
reg:pseudohubererror
huber_slope
1.0
n_estimators
300
Appendix
Table 13: Fixed hyperparameters used by the tree baselines.
Sparse feature selection is critical for high-dimensional machine learning, yet traditional ℓ1-regularized methods are often brittle under observational noise and spurious correlations, leading to unstable feature supports and degraded generalization. Although adversarial training has been widely used to improve model robustness, its interaction with hierarchical sparse feature selection remains underexplored. In this work, we propose Adversarial LassoNet (AdLNet), a stability-driven sparse feature selection framework that integrates input-space adversarial perturbations with the hierarchical sparsity mechanism of LassoNet. We derive a tractable first-order adversarial approximation under local smoothness assumptions and provide an NTK-inspired spectral analysis to characterize how perturbation-driven training can reduce gradient concentration. Experiments on high-dimensional SERS data, six public benchmark datasets, and ColoredMNIST show that AdLNet maintains competitive sparse-selection performance while improving out-of-distribution robustness by 4.4% and feature support reproducibility by 6.3% under nearly matched support sparsity on ColoredMNIST. On the high-dimensional lung cancer screening dataset, AdLNet achieves a 5.3% test accuracy gain and a 6.0% AUC improvement over vanilla LassoNet. Code and dataset are available at https://github.com/719573/Adversarial-LassoNet.
Zhen Huang, Peicheng Xu, Junbiao Pang +1
Beijing University of Technology · The First Affiliated Hospital of Zhejiang University School of Medicine
Most real-world datasets used for training supervised learning models are contaminated with noisy data and outliers leading to large prediction errors. This paper proposes a new approach for achieving robustness where the learning rate is modulated by a factor that is sensitive to outliers. In this approach a reduction of the learning rate is shown to be achieved by using alternate loss functions that are infinitely differentiable, strictly convex or quasiconvex and more closely approximate the absolute error than Huber and log-cosh losses. A comparison of the performance of regression models trained with different loss functions on a wide variety of benchmarks and datasets is presented to demonstrate the superior performance of the Square Root Loss (SRL) and Smooth Mean Absolute Error (SMAE) losses proposed in this paper. Two new robust linear regression models are presented. Highly vectorized robust parameter update formulae that take advantage of modern GPUs for both stochastic and batch gradient descent are presented.
Mathew Mithra Noel, Arindam Banerjee, Yug D. Oswal +2
aSchool of Electrical Engineering, Vellore Institute of Technology, Vellore, 632 014, Tamil Nadu, India · bErnst and Young GDS, Godrej Waterside, Ring Rd, DP Block, Sector V, Bidhannagar, Kolkata, 700 091, West Bengal, India · cSchool of Computer Science and Engineering, Vellore Institute of Technology, Vellore, 632 014, Tamil Nadu, India
Broad Learning System (BLS) offers an efficient alternative to deep architectures by enabling fast learning through randomized feature mapping and closed-form solutions. However, its reliance on squared error loss makes it highly sensitive to noise, outliers, and corrupted labels, limiting its reliability in real-world scenarios. To address this limitation, we propose Wave-BLS, a robust broad learning framework that integrates the wave loss function, which is asymmetric, bounded, and smooth, enabling controlled penalization of large errors. The proposed formulation replaces the standard least-squares objective with a wave-loss-based optimization problem, solved efficiently using a Nesterov accelerated gradient (NAG)-based scheme without requiring matrix inversion, thereby improving scalability. Extensive experiments on 30 UCI benchmark datasets demonstrate that Wave-BLS consistently outperforms classical BLS and several robust variants. Statistical validation using Friedman and Nemenyi post-hoc tests confirms the significance of the observed improvements. Furthermore, robustness evaluations under controlled noise and outlier injection reveal that Wave-BLS exhibits substantially slower performance degradation compared to BLS, even in challenging contamination settings. These results establish Wave-BLS as a stable and robust alternative to existing broad learning models for learning under data uncertainty.
Mushir Akhtar, A. Varshney, A. Quadir +3
Department of Mathematics Indian Institute of Technology Indore