Parameter-Efficient Distributionally Robust Adaptation of Tabular Foundation Models under Subpopulation Shift
Organizations: Pohang University of Science and Technology
Abstract
Despite strong mean accuracy, tabular foundation models (TFMs) can perform poorly on underrepresented groups under subpopulation shift, where group proportions change between training and deployment. We propose DR-TFM, a parameter-efficient distributionally robust adaptation framework that requires no true group annotations. DR-TFM adjusts attention to labeled context examples by fine-tuning an existing query scaling network or adding and training one, while keeping all other parameters fixed. We instantiate the framework with two robust objectives using estimated groups or source conditional distributions derived from training data. For TabPFN-3, adaptation updates only 0.016% of the pretrained model's parameters. Across five tabular benchmarks, DR-TFM achieves substantially higher average worst-group accuracy than pretrained TFMs and the compared robust baselines without true group annotations, while maintaining competitive mean group accuracy. DR-TFM also improves average worst-group accuracy on ACS Income and across four additional TFMs.
Figures & tables
| Methods | Adult | Bank | Default | Shoppers | Taxi | Avg. | ||||||
| Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | |
| Standard methods | ||||||||||||
| ERM-MLP | 73.45 | 45.29 | 71.86 | 42.03 | 63.70 | 28.52 | 76.44 | 50.11 | 75.27 | 57.37 | 72.14 | 44.67 |
| XGBoost | 77.09 | 52.58 | 68.69 | 33.87 | 64.78 | 30.72 | 78.03 | 55.42 | 75.81 | 57.57 | 72.88 | 46.03 |
| CatBoost | 76.50 | 51.33 | 68.94 | 34.43 | 64.75 | 31.03 | 77.50 | 54.24 | 76.04 | 58.71 | 72.75 | 45.95 |
| TFMs | ||||||||||||
| Methods | AZ | MA | MI | Avg. | AZ MA | MA MI | MI AZ | Avg. | ||||||||
| Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | |
| Standard methods | ||||||||||||||||
| ERM-MLP | 77.20 | 59.42 | 80.82 | 73.82 | 75.69 | 57.99 | 77.90 | 63.74 | 76.82 | 60.23 | 79.13 | 71.34 | 75.56 | 56.31 | 77.17 | 62.63 |
| XGBoost | 77.64 | 58.99 | 82.26 | 76.38 | 77.50 | 62.77 | 79.13 | 66.05 | 77.96 | 61.99 | 79.67 | 72.01 | 76.82 | 58.99 | 78.15 | 64.33 |
| CatBoost | 77.35 | 58.26 | 81.97 | 75.59 | 77.27 | 62.43 | 78.86 | 65.43 | 77.66 | 61.15 | 79.60 | 72.31 | 76.68 | 58.71 | 77.98 | 64.06 |
| TFMs | ||||||||||||||||
| GEORGE | LSR | MixturePFN | BETA | TabPFN | TabPFN (ft) | DR-TFM (GR) | DR-TFM (MS) | |
|---|---|---|---|---|---|---|---|---|
| Seconds | 77.0 | 5,156.4 | 229.5 | 1,020.6 | 5.4 | 214.1 | 23.8 | 30.6 |
| Memory (GB) | 1.21 | 3.12 | 2.17 | 14.33 | 0.61 | 9.25 | 1.07 | 0.98 |
| Fine-tuned parameters | Trainable (%) | ERM (%) | DRO (%) | Seconds | Memory (GB) |
|---|---|---|---|---|---|
| Input adapter | 0.063–0.155 | 45.45 | 66.64 (+21.19) | 43.5 | 9.80 |
| Encoder | 99.196 | 48.39 | 68.92 (+20.53) | 241.2 | 10.22 |
| Full backbone | 48.00 | 68.98 (+20.98) | 246.6 | 10.31 | |
| Decoder | 0.804 | 48.68 | 71.93 (+23.25) | 24.0 | 1.10 |
| Query scaling | 0.016 | 47.34 | 72.57 (+25.23) | 23.8 | 1.07 |
| Method | Five tabular benchmarks (Avg.) | ACS within-state (Avg.) | ACS transfer (Avg.) | |||
|---|---|---|---|---|---|---|
| Mean | Worst | Mean | Worst | Mean | Worst | |
| TabPFN | 73.75 | 47.94 | 79.64 | 66.80 | 78.15 | 64.17 |
| + Ours (0.016%) | 79.31 (+5.56) | 72.57 (+24.63) | 79.21 (-0.43) | 71.06 (+4.26) | 78.54 (+0.39) | 68.32 (+4.15) |
| TabPFN-3.5 | 74.39 | 49.26 | 79.79 | 67.10 | 78.25 | 64.34 |
| + Ours (0.004%) | 79.42 (+5.03) | 71.47 (+22.21) | 79.88 (+0.09) | 72.48 (+5.38) | 79.19 (+0.94) | 71.51 (+7.17) |
| EXAONE | 73.72 | 47.68 | 79.52 | 66.16 | 78.27 | 64.19 |
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Train | Validation | Test | Total |
|---|---|---|---|---|
| Adult | 26,048 | 6,513 | 16,281 | 48,842 |
| Bank | 32,551 | 8,138 | 4,522 | 45,211 |
| Default | 21,600 | 5,400 | 3,000 | 30,000 |
| Shoppers | 8,877 | 2,220 | 1,233 | 12,330 |
| Taxi | 9,139 | 2,285 | 1,270 | 12,694 |
| Setting | Train | Validation | Test | Total |
|---|---|---|---|---|
| AZ | 23,959 | 5,990 | 3,328 | 33,277 |
| MA | 28,881 | 7,221 | 4,012 | 40,114 |
| MI | 36,005 | 9,002 | 5,001 | 50,008 |
| AZ MA | 26,621 | 6,656 | 40,114 | 73,391 |
| MA MI | 32,091 | 8,023 | 50,008 | 90,122 |
| MI AZ | 40,006 | 10,002 | 33,277 | 83,285 |
| Method | Training | Epochs / steps | Selection criterion | Learning rate schedule |
|---|---|---|---|---|
| ERM-MLP | ERM, from scratch | 1000 | validation accuracy | multiply by after epochs without improvement in mean training loss |
| GroupDRO | GroupDRO over the true groups, from scratch | 1000 | validation worst-group accuracy (true groups) | same as above |
| GEORGE stage 1 | ERM, from scratch (feature extractor) | 1000 | validation accuracy | none |
| GEORGE stage 2 | GroupDRO over the estimated clusters, newly initialized model | 300 | worst-group accuracy over the estimated validation clusters | none |
| BPA | base ERM, then retraining with cluster reweighting | 1000 / 300 | worst-group accuracy over the estimated validation clusters | none |
| SPARE | base ERM, then retraining with importance sampling | 1000 / 300 | worst-group accuracy over the estimated validation clusters | none |
| Validation | 70.43 (1.21) | 70.44 (1.18) | 70.44 (1.13) | 70.38 (1.15) | 70.37 (1.15) | 70.39 (1.12) |
|---|
| Method | LR | Batch size | Budget | Patience | Weight decay | Selected quantities | ||
| DR-TFM (GR; -means) | 0.1 | 5 | 0.01 | full | 1000 | 5 | 0.005 | per class by silhouette over |
| DR-TFM (GR; GMM) | 0.1 | 5 | 0.01 | full | 1000 | 5 | 0.005 | per class by silhouette over |
| DR-TFM (GR) ∗ | 0.1 | 5 | 0.01 | full | 1000 | 5 | 0.005 | none (groups given) |
| DR-TFM (MS) | — | — | 0.001 | 1024 | 300 | 5 | once on Adult; fixed thereafter | |
| DR-TFM (MS) ∗ | — | — | 0.001 | 1024 | 300 | 5 | none (groups given) | |
| ERM-MLP | — | — | 0.001 | 4096 | 1000 | 20 | 0 | none |
| Method | Trained component | Settings |
|---|---|---|
| TabPFN (fine-tuned) | Full backbone | Learning rate , weight decay , 30 epochs, patience 8 |
| TabPFN (class-balanced) | Full backbone | As TabPFN (fine-tuned), with inverse-class-frequency weights in the cross-entropy |
| BETA | Input adapter | 16 ensemble members, learning rate , weight decay , batch size 1024, 30 epochs, patience 5 |
| MixturePFN | Expert models | Context size 3000; experts use the TabPFN fine-tuning settings above |
| LoCalPFN | None | Retrieval context of 1000 neighbors per query batch |
| TuneTables | Soft prompt | 10 tokens, epochs, patience 5, TabPFN v1 |
| Method | Range (pp) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | |
| DR-TFM (GR; -means) | 79.31 | 72.57 | 77.42 | 68.37 | 75.97 | 63.90 | 75.09 | 61.39 | 4.22 | 11.18 |
| (0.17) | (0.64) | (0.33) | (1.29) | (0.63) | (3.00) | (0.77) | (1.12) | |||
| GEORGE ( -means) | 73.79 | 63.80 | 61.80 | 43.96 | 64.22 | 48.25 | 58.64 | 44.45 | 15.15 | 19.84 |
| (0.48) | (1.21) | (2.28) | (4.56) | (1.04) | (2.78) | (5.89) | (9.45) | |||
| (pp) | +5.52 | +8.77 | +15.62 | +24.41 | +11.75 | +15.65 | +16.45 | +16.94 | 10.93 | 15.64 |
| Method | Range (pp) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | |
| DR-TFM (GR; -means) | 79.21 | 71.06 | 80.00 | 72.99 | 79.53 | 72.22 | 79.70 | 72.59 | 0.78 | 1.93 |
| (0.14) | (0.12) | (0.12) | (0.40) | (0.38) | (1.11) | (0.04) | (0.39) | |||
| GEORGE ( -means) | 74.36 | 66.15 | 64.70 | 54.67 | 65.41 | 56.41 | 66.75 | 46.94 | 9.66 | 19.22 |
| (3.03) | (2.31) | (1.35) | (4.89) | (2.71) | (1.27) | (6.77) | (9.64) | |||
| (pp) | +4.85 | +4.91 | +15.29 | +18.32 | +14.11 | +15.81 | +12.95 | +25.66 | 10.44 | 20.75 |
| Methods | Adult | Bank | Default | Shoppers | Taxi | Avg. | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | |
| Standard methods | ||||||||||||
| ERM-MLP | 73.45 | 45.29 | 71.86 | 42.03 | 63.70 | 28.52 | 76.44 | 50.11 | 75.27 | 57.37 | 72.14 | 44.67 |
| (0.16) | (0.90) | (1.80) | (4.18) | (0.48) | (0.21) | (0.47) | (1.85) | (0.60) | (1.44) | (0.70) | (1.72) | |
| XGBoost | 77.09 | 52.58 | 68.69 | 33.87 | 64.78 | 30.72 | 78.03 | 55.42 | 75.81 | 57.57 | 72.88 | 46.03 |
| (0.17) | (0.48) | (0.31) | (0.92) | (0.11) | (0.13) | (0.19) | (0.24) | (0.21) | (0.16) | (0.20) | (0.39) | |
| Methods | Adult | Bank | Default | Shoppers | Taxi | Avg. | ||||||
| Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | |
| Robust methods | ||||||||||||
| Fair-TabICL (group-balanced) | 80.11 | 72.34 | 86.58 | 78.54 | 69.86 | 57.73 | 86.24 | 83.38 | 77.45 | 68.03 | 80.05 | 72.00 |
| (0.23) | (0.91) | (0.09) | (0.10) | (0.23) | (0.71) | (0.85) | (0.91) | (0.32) | (0.49) | (0.34) | (0.62) | |
| Fair-TabICL (uncertainty) | 61.80 | 21.35 | 71.52 | 40.51 | 64.93 | 31.18 | 77.80 | 52.81 | 75.94 | 57.61 | 70.40 | 40.69 |
| (0.74) | (1.65) | (0.14) | (0.32) | (0.12) | (0.13) | (0.43) | (1.57) | (0.14) | (0.52) | (0.31) | (0.84) | |
| Methods | AZ | MA | MI | Avg. | AZ MA | MA MI | MI AZ | Avg. | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | |
| Standard methods | ||||||||||||||||
| ERM-MLP | 77.20 | 59.42 | 80.82 | 73.82 | 75.69 | 57.99 | 77.90 | 63.74 | 76.82 | 60.23 | 79.13 | 71.34 | 75.56 | 56.31 | 77.17 | 62.63 |
| (0.54) | (1.59) | (0.07) | (0.61) | (0.32) | (1.88) | (0.31) | (1.36) | (0.62) | (2.91) | (0.09) | (0.66) | (0.30) | (2.40) | (0.34) | (1.99) | |
| XGBoost | 77.64 | 58.99 | 82.26 | 76.38 | 77.50 | 62.77 | 79.13 | 66.05 | 77.96 | 61.99 | 79.67 | 72.01 | 76.82 | 58.99 | 78.15 | 64.33 |
| (0.27) | (0.76) | (0.25) | (0.12) | (0.11) | (0.40) | (0.21) | (0.43) | (0.14) | (0.36) | (0.06) | (0.13) | (0.10) | (0.30) | (0.10) | (0.27) | |
| Methods | AZ | MA | MI | Avg. | AZ MA | MA MI | MI AZ | Avg. | ||||||||
| Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | |
| Robust methods | ||||||||||||||||
| Fair-TabICL (group-balanced) | 79.43 | 74.38 | 81.42 | 78.68 | 78.96 | 75.19 | 79.94 | 76.09 | 80.10 | 78.18 | 79.56 | 75.01 | 78.87 | 75.41 | 79.51 | 76.20 |
| (0.24) | (0.09) | (0.03) | (0.82) | (0.20) | (1.21) | (0.16) | (0.71) | (0.08) | (0.03) | (0.06) | (0.54) | (0.15) | (1.33) | (0.10) | (0.63) | |
| Fair-TabICL (uncertainty) | 77.72 | 58.40 | 82.36 | 77.46 | 76.71 | 62.07 | 78.93 | 65.98 | 77.41 | 60.40 | 79.59 | 73.40 | 76.24 | 57.21 | 77.75 | 63.67 |
| (0.17) | (0.57) | (0.05) | (0.21) | (0.31) | (0.61) | (0.18) | (0.47) | (0.08) | (0.39) | (0.12) | (0.25) | (0.13) | (0.13) | (0.11) | (0.26) | |
| Method | Adult | Bank | Default | Shoppers | Taxi | Avg. |
|---|---|---|---|---|---|---|
| XGBoost | 5.5 | 6.4 | 7.1 | 4.6 | 4.1 | 5.5 |
| CatBoost | 22.8 | 25.2 | 24.1 | 12.8 | 12.8 | 19.5 |
| TabPFN | 7.5 | 7.8 | 5.5 | 3.2 | 3.1 | 5.4 |
| ERM-MLP | 32.0 | 36.6 | 29.9 | 15.1 | 14.8 | 25.7 |
| DR-TFM (GR; -means) | 25.6 | 50.2 | 24.6 | 9.1 | 9.4 | 23.8 |
| DR-TFM (MS; -means) | 35.9 | 45.0 | 36.6 | 16.9 | 18.8 | 30.6 |
| Method | Adult | Bank | Default | Shoppers | Taxi | Avg. |
|---|---|---|---|---|---|---|
| TabPFN | 0.77 | 0.77 | 0.62 | 0.44 | 0.43 | 0.61 |
| DR-TFM (GR; -means) | 1.37 | 1.62 | 1.15 | 0.60 | 0.61 | 1.07 |
| DR-TFM (MS; -means) | 1.27 | 1.41 | 1.02 | 0.61 | 0.58 | 0.98 |
| GEORGE ( -means) | 1.62 | 1.63 | 1.60 | 0.71 | 0.51 | 1.21 |
| TabPFN (fine-tuned) | 10.98 | 14.45 | 11.81 | 5.25 | 3.75 | 9.25 |
| LSR (all stages) | 3.56 | 5.07 | 4.82 | 1.50 | 0.67 | 3.12 |
| Adaptation | Adult | Bank | Default | Shoppers | Taxi | Avg. | ||||||
| Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | |
| Input adapter | ||||||||||||
| ERM | 74.34 | 46.89 | 69.39 | 36.98 | 64.39 | 30.34 | 78.31 | 55.67 | 75.38 | 57.36 | 72.36 | 45.45 |
| (0.67) | (1.58) | (1.53) | (2.81) | (0.10) | (0.53) | (0.91) | (2.12) | (0.14) | (0.22) | (0.67) | (1.45) | |
| DRO | 78.83 | 65.90 | 82.70 | 73.21 | 67.92 | 47.05 | 84.96 | 79.60 | 77.23 | 67.43 | 78.33 | 66.64 |
| (0.51) | (1.58) | (1.34) | (1.38) | (0.38) | (1.44) | (0.61) | (0.57) | (0.27) | (0.97) | (0.62) | (1.19) | |
| Fine-tuned parameters | Adult | Bank | Default | Shoppers | Taxi | Avg. |
|---|---|---|---|---|---|---|
| Input adapter | 56.5 | 63.0 | 40.8 | 26.8 | 30.3 | 43.5 |
| Encoder | 214.2 | 439.1 | 364.7 | 80.0 | 107.8 | 241.2 |
| Full backbone | 213.8 | 438.3 | 383.8 | 90.3 | 106.8 | 246.6 |
| Decoder | 29.9 | 45.3 | 25.9 | 9.6 | 9.5 | 24.0 |
| Query scaling | 25.6 | 50.2 | 24.6 | 9.1 | 9.4 | 23.8 |
| Fine-tuned parameters | Adult | Bank | Default | Shoppers | Taxi | Avg. |
|---|---|---|---|---|---|---|
| Input adapter | 12.58 | 15.45 | 10.86 | 5.02 | 5.09 | 9.80 |
| Encoder | 12.22 | 16.10 | 13.20 | 5.32 | 4.27 | 10.22 |
| Full backbone | 12.34 | 16.24 | 13.29 | 5.36 | 4.31 | 10.31 |
| Decoder | 1.38 | 1.68 | 1.18 | 0.62 | 0.63 | 1.10 |
| Query scaling | 1.37 | 1.62 | 1.15 | 0.60 | 0.61 | 1.07 |
| Method | Five tabular benchmarks | ACS Income | ||
|---|---|---|---|---|
| Mean | Worst | Mean | Worst | |
| ERM-MLP | 72.14 | 44.67 | 77.54 | 63.19 |
| -DRO † | 70.00 | 47.82 | 74.62 | 62.95 |
| TabPFN- -DRO ( ) | 73.34 | 47.34 | 78.73 | 65.26 |
| TabPFN- -DRO ( ) | 73.16 | 47.19 | 78.02 | 63.27 |
| TabPFN- -DRO ( ) | 69.65 | 37.75 | 76.93 | 59.40 |
| Methods | Adult | Bank | Default | Shoppers | Taxi | Avg. | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | |
| TabPFN | 75.33 | 49.17 | 73.07 | 42.96 | 64.77 | 31.09 | 79.59 | 58.29 | 76.01 | 58.19 | 73.75 | 47.94 |
| (0.05) | (0.28) | (0.31) | (0.76) | (0.04) | (0.04) | (0.32) | (0.52) | (0.25) | (0.53) | (0.19) | (0.43) | |
| + Ours | 79.64 | 67.90 | 85.01 | 80.14 | 69.42 | 60.37 | 85.79 | 81.66 | 76.69 | 72.79 | 79.31 | 72.57 |
| (0.20) | (1.61) | (0.36) | (0.53) | (0.62) | (1.70) | (0.32) | (0.50) | (0.19) | (0.62) | (0.34) | (0.99) | |
| (+4.31) | (+18.73) | (+11.94) | (+37.18) | (+4.65) | (+29.28) | (+6.20) | (+23.37) | (+0.68) | (+14.60) | (+5.56) | (+24.63) | |
| Methods | AZ | MA | MI | Avg. | AZ MA | MA MI | MI AZ | Avg. | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | Mean | Worst | |
| TabPFN | 78.19 | 59.87 | 83.01 | 77.02 | 77.72 | 63.50 | 79.64 | 66.80 | 77.95 | 61.57 | 79.69 | 71.86 | 76.81 | 59.08 | 78.15 | 64.17 |
| (0.32) | (0.50) | (0.12) | (0.27) | (0.03) | (0.20) | (0.16) | (0.32) | (0.07) | (0.18) | (0.08) | (0.25) | (0.05) | (0.12) | (0.07) | (0.18) | |
| + Ours | 78.81 | 70.36 | 79.96 | 70.93 | 78.87 | 71.90 | 79.21 | 71.06 | 78.43 | 69.70 | 78.73 | 66.43 | 78.46 | 68.84 | 78.54 | 68.32 |
| (0.13) | (1.00) | (0.59) | (1.31) | (0.30) | (0.86) | (0.34) | (1.06) | (0.15) | (0.72) | (0.68) | (0.34) | (0.15) | (0.59) | (0.32) | (0.55) | |
| (+0.62) | (+10.49) | (-3.05) | (-6.09) | (+1.15) | (+8.40) | (-0.43) | (+4.26) | (+0.48) | (+8.13) | (-0.96) | (-5.43) | (+1.65) | (+9.76) | (+0.39) | (+4.15) | |