Sparsifying Stochasticity, Not Capacity: Partial Stochasticity via Deep Weight Factorization of Prior Scales
Organizations: Université Grenoble Alpes, France · KAUST, Saudi Arabia
Abstract
Bayesian neural networks need not be fully stochastic to be universal conditional density approximators, but it remains open which parameters should be stochastic. We learn this split by applying deep weight factorization to the prior scales, which are the standard deviations of the parameter priors, while fitting the functional prior to a Gaussian process with a maximum mean discrepancy objective. A parameter whose prior scale falls below a cutoff becomes deterministic and is optimized during inference, so the regularizer sparsifies stochasticity rather than capacity. We give a certificate for universal conditional density approximation that is checkable in linear time, together with a minimal repair when it fails. We further show that the common hybrid scheme of sampling some parameters and optimizing the others is stochastic approximation for a type-II maximum a posteriori objective, and that coupled step sizes can leave a tracking error that does not vanish as the step size shrinks. On a bimodal target, the learned split stays close to an unconstrained reference across all budgets and is insensitive to the cutoff, while random masks that distribute the same prior scales across layers are worse by up to two orders of magnitude. On UCI benchmarks, our method performs on par with a fully stochastic network while keeping about half of its parameters deterministic.
Figures & tables
| variant | condition violated | |
| sampling floor | — | .020 |
| point predictor | — | .681 |
| noise input, | reference, none | |
| layer 1, | none | |
| layer 1, , weight | vanishes at | |
| output layer, | bias coverage |
| rmse ( ) | .316 .033 | .296 .029 | .291 .027 |
|---|---|---|---|
| nll ( ) | .274 .095 | .205 .074 | .193 .066 |
| ece ( ) | .085 .059 | .115 .083 | .119 .074 |
| .04 .01 | 23.97 .48 | 76.94 4.41 |
| Classification ( ) | |||
|---|---|---|---|
| Method | acc ( ) | nll ( ) | ece ( ) |
| DWF-PSNN | 2.38 | 2.63 | 2.88 |
| ALS-BNN | 2.63 | 2.63 | 3.13 |
| FLS-BNN | 2.88 | 2.88 | 2.50 |
| RAND-PSNN | 2.38 | 1.88 | 2.63 |
| SNI-PSNN | 4.75 | 5.00 | 3.88 |
| Dataset | learned | repair | random | |
|---|---|---|---|---|
| BANKNOTE | 4 | 1.00 | 0.0 | 1.00 |
| HTRU2 | 8 | 1.00 | 0.0 | 1.00 |
| CONCRETE | 8 | 0.90 | 0.1 | 0.97 |
| WINE | 10 | 0.00 | 5.5 | 0.09 |
| BIKESHARING | 12 | 0.00 | 7.3 | 0.01 |
| BOSTON | 13 | 0.00 | 7.8 | 0.00 |
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Sample size | Parameters | Citation | |
|---|---|---|---|---|
| CONCRETE | (8, 1) | 1030 | 11101 | ( Yeh, 1998 ) |
| BOSTON | (13, 1) | 506 | 11601 | ( Harrison and Rubinfeld, 1978 ) |
| BIKESHARING | (12, 1) | 731 | 11501 | ( Fanaee-T, 2013 ) |
| WINE | (10, 1) | 4898 | 11301 | ( Cortez et al., 2009 ) |
| ONLINESHOPPERS | (17, 2) | 12330 | 12102 | ( Sakar and Kastro, 2018 ) |
| CREDIT | (23, 2) | 30000 | 12702 | ( Yeh, 2009 ) |
| acc in % | nll | ece | Deterministic parameters in % | |
|---|---|---|---|---|
| DWF-PSNN | 100.00 0.00 | 0.009 0.002 | 0.009 0.002 | 55.32 2.56 |
| ALS-BNN | 100.00 0.00 | 0.004 0.001 | 0.004 0.000 | 0.00 |
| FLS-BNN | 100.00 0.00 | 0.004 0.001 | 0.004 0.001 | 95.37 |
| RAND-PSNN | 100.00 0.00 | 0.005 0.002 | 0.005 0.001 | 55.32 2.56 |
| SNI-PSNN | 99.20 0.63 | 0.032 0.010 | 0.025 0.007 | 90.00 |
| acc in % | nll | ece | Deterministic parameters in % | |
|---|---|---|---|---|
| DWF-PSNN | 97.83 0.48 | 0.069 0.016 | 0.007 0.003 | 57.31 2.79 |
| ALS-BNN | 97.59 0.50 | 0.078 0.016 | 0.008 0.003 | 0.00 |
| FLS-BNN | 97.56 0.56 | 0.079 0.015 | 0.008 0.003 | 91.97 |
| RAND-PSNN | 97.82 0.49 | 0.070 0.015 | 0.007 0.003 | 57.31 2.79 |
| SNI-PSNN | 97.41 0.58 | 0.086 0.016 | 0.009 0.002 | 90.00 |
| acc in % | nll | ece | Deterministic parameters in % | |
|---|---|---|---|---|
| DWF-PSNN | 89.43 0.64 | 0.243 0.020 | 0.016 0.005 | 61.01 3.43 |
| ALS-BNN | 89.25 0.74 | 0.261 0.019 | 0.024 0.009 | 0.00 |
| FLS-BNN | 89.25 0.85 | 0.256 0.020 | 0.018 0.006 | 85.13 |
| RAND-PSNN | 89.48 0.69 | 0.243 0.020 | 0.015 0.005 | 60.91 3.44 |
| SNI-PSNN | 88.70 0.90 | 0.292 0.017 | 0.024 0.012 | 90.00 |
| acc in % | nll | ece | Deterministic parameters in % | |
|---|---|---|---|---|
| DWF-PSNN | 78.02 0.61 | 0.508 0.010 | 0.022 0.010 | 60.90 2.61 |
| ALS-BNN | 78.03 0.61 | 0.502 0.009 | 0.019 0.006 | 0.00 |
| FLS-BNN | 78.03 0.61 | 0.506 0.009 | 0.015 0.005 | 81.11 |
| RAND-PSNN | 78.02 0.61 | 0.500 0.009 | 0.023 0.006 | 60.68 2.61 |
| SNI-PSNN | 78.02 0.61 | 0.527 0.007 | 0.014 0.013 | 90.00 |
| rmse | nll | ece | Deterministic parameters in % | |
|---|---|---|---|---|
| DWF-PSNN | 0.298 0.027 | 0.214 0.073 | 0.107 0.066 | 48.69 1.27 |
| ALS-BNN | 0.300 0.032 | 0.283 0.068 | 0.134 0.062 | 0.00 |
| FLS-BNN | 0.335 0.030 | 0.286 0.085 | 0.085 0.062 | 91.89 |
| RAND-PSNN | 0.331 0.034 | 0.322 0.113 | 0.075 0.073 | 48.69 1.27 |
| SNI-PSNN | 0.332 0.039 | 0.326 0.133 | 0.097 0.088 | 90.00 |
| rmse | nll | ece | Deterministic parameters in % | |
|---|---|---|---|---|
| DWF-PSNN | 0.771 0.023 | 2.210 0.145 | 0.127 0.015 | 47.12 2.93 |
| ALS-BNN | 0.723 0.031 | 1.049 0.065 | 0.017 0.015 | 0.00 |
| FLS-BNN | 0.856 0.024 | 3.395 0.202 | 0.207 0.017 | 90.27 |
| RAND-PSNN | 0.841 0.035 | 3.314 0.294 | 0.149 0.050 | 47.07 2.93 |
| SNI-PSNN | 0.844 0.036 | 3.338 0.304 | 0.135 0.029 | 90.00 |
| rmse | nll | ece | Deterministic parameters in % | |
|---|---|---|---|---|
| DWF-PSNN | 0.105 0.034 | -0.159 0.034 | 0.585 0.196 | 47.81 2.95 |
| ALS-BNN | 0.155 0.037 | 0.049 0.038 | 0.432 0.084 | 0.00 |
| FLS-BNN | 0.101 0.040 | -0.164 0.041 | 0.614 0.213 | 88.70 |
| RAND-PSNN | 0.101 0.037 | -0.174 0.050 | 0.600 0.191 | 47.75 2.95 |
| SNI-PSNN | 0.104 0.033 | -0.173 0.033 | 0.674 0.180 | 90.00 |
| rmse | nll | ece | Deterministic parameters in % | |
|---|---|---|---|---|
| DWF-PSNN | 0.386 0.113 | 0.482 0.331 | 0.209 0.145 | 48.03 2.53 |
| ALS-BNN | 0.451 0.107 | 0.470 0.170 | 0.226 0.166 | 0.00 |
| FLS-BNN | 0.392 0.116 | 0.513 0.337 | 0.213 0.172 | 87.93 |
| RAND-PSNN | 0.398 0.102 | 0.608 0.385 | 0.196 0.193 | 47.97 2.53 |
| SNI-PSNN | 0.408 0.106 | 0.652 0.425 | 0.310 0.305 | 90.00 |
| Deterministic parameters | Stochastic | Deterministic | |||
|---|---|---|---|---|---|
| Dataset | Layer 1 | Layer 2 | Output layer | first-layer biases | output bias |
| ONLINESHOPPERS | 9–17 | 6/10 | |||
| CREDIT | 7–20 | 10/10 | |||
| BANKNOTE | 8–18 | 0/10 | |||
| HTRU2 | 10–18 | 7/10 | |||
| CONCRETE | 11–22 | 4/10 | |||
| variant | condition violated | fixed | learned | |
|---|---|---|---|---|
| sampling floor | — | — | 0.020 | 0.020 |
| point predictor | — | 0 | 0.681 | 0.681 |
| noise input, | reference, none | 0 | ||
| first layer, | none | 1 | ||
| first layer, , weight | vanishes at | 1 | ||
| Hyperparameter | Value |
|---|---|
| Batch size | 64 |
| Number of iterations | 10000 |
| Regularization ( ) | (9, 12, 15) / D |
| Number of function samples | 100 |
| Learning rate | 0.01 |
| Kernel | Polynomial of degree 2 |
| Hyperparameter | Value |
| Batch size | 64 |
| Number of samples per chain | 30 |
| Discarded samples before sampling | 10 |
| Thinning interval | 200 |
| Number of chains | 4 |
| Momentum decay | 0.02 |
| Hyperparameter | Value |
|---|---|
| Batch size | 64 |
| Prior precision | 1.0 |
| Stochastic fraction | 0.1 |
| Noise variance | 0.1 |
| Epochs | 1000 |
| Learning rate | 0.01 |
| Hyperparameter | Value |
| Batch size | 64 |
| Number of samples per chain | 30 |
| Discarded samples before sampling | 10 |
| Thinning interval | 200 |
| Number of chains | 4 |
| Momentum decay | 0.02 |
| Hyperparameter | Value |
|---|---|
| Batch size | 64 |
| Prior precision | 1.0 |
| Stochastic fraction | 0.1 |
| Noise variance | 0.1 |
| Epochs | 1000 |
| Learning rate | 0.01 |
| Hyperparameter | Value |
|---|---|
| Batch size | 128 |
| Number of iterations | 15000 |
| Regularization ( ) | (1.2, 2.4, 4.8) / D |
| Number of function samples | 100 |
| Learning rate | 0.03 |
| Kernel | Polynomial of degree 2 |
| Hyperparameter | Value |
| Batch size | 64 |
| Number of samples per chain | 60 |
| Discarded samples before sampling | 20 |
| Thinning interval | 100 |
| Number of chains | 4 |
| Momentum decay | 0.01 |
| Hyperparameter | Value |
|---|---|
| Batch size | 64 |
| Number of samples per chain | 100 |
| Discarded samples before sampling | 10 |
| Number of burn-in steps | 10000 |
| Thinning interval | 20 |
| Number of chains | 4 |