Beyond Correctness: Evaluating Semantic Knowledge in Cross-Table Transfer
Organizations: Korea University
Abstract
Semantic knowledge is increasingly used to bridge heterogeneous schemas in tabular learning, but how much does that knowledge actually improve prediction? Studies in tabular learning commonly answer this question through semantic ablations that modify or suppress the supplied semantic knowledge. We show that these ablations can lead to misleading conclusions about predictive benefit: poor performance under altered semantics may be taken as evidence that the intended knowledge is beneficial. Across real and controlled experiments, altering semantic content can produce large performance differences even when the model gains little predictive benefit from having that semantic knowledge in the first place. To separate these effects, we distinguish two quantities: content sensitivity and predictive utility. Content sensitivity measures the change in performance when semantic content is altered, whereas predictive utility measures the benefit of the intended semantic knowledge relative to a suitable reference without that knowledge. This distinction motivates an evaluation framework in which the control is chosen according to the question being asked: altered controls assess sensitivity to semantic content, whereas claims that semantic knowledge improves prediction require a suitable reference. Even then, predictive utility is not fixed; it varies across suitable references and decreases when the reference can more easily recover the tested knowledge from other inputs or labeled examples. In a bounded audit of 25 semantic-ablation comparisons across nine studies, only one of 18 explicit predictive-utility claims is paired with a control that clearly isolates the tested semantic contribution. Together, these findings motivate a simple evaluation principle: semantic-ablation controls should be chosen and interpreted according to the question they are intended to answer.
Figures & tables
| Knowledge | Intended | Altered | Reference |
|---|---|---|---|
| Measurement | Specified coding or numerical conversion | Altered coding or numerical conversion | Raw source values |
| Correspondence | Specified source–target pairing | Permuted source–target pairing | No shared pairing; separate source and target columns |
| Relation | Specified relation and direction | Reversed direction or random relation | No added value or constraint |
| Setting | Data and task | Primary model | Metric |
|---|---|---|---|
| TabLLM reanalysis | Nine original-study datasets; released-code reanalysis with the name-masked reference | TabLLM (released) | AUROC |
| Clinical | NH–KN and pooled NH/KN–HRS; binary questionnaire and continuous examination endpoints | XGB | AUROC (binary); negative SD-normalized RMSE (continuous) |
| Hydrological | CAMELS-US to CAMELS-GB v2; 120 basins (CAMELS-120) | XGB | Negative log-RMSE |
| Controlled benchmark | Eight-variable NHANES panel with observed values and missingness retained; constructed schemas and synthetic targets | XGB | Negative nMSE |
| Setting | Altered condition | Content sensitivity | Predictive utility | Reference Altered |
|---|---|---|---|---|
| TabLLM feature-name ablation | ||||
| 9-dataset mean | List Permuted Names | |||
| Credit-g | List Permuted Names | |||
| Bank | List Permuted Names | |||
| Real cross-table transfer: relation knowledge | ||||
| NH–KN | Reversed | |||
| Claim type | Supported | Unsupported | Indeterminate |
|---|---|---|---|
| Content sensitivity ( ) | 3 | 1 | 0 |
| Predictive utility ( ) | 1 | 11 | 6 |
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Target setup | Model | Model settings | Runs | Inference unit |
|---|---|---|---|---|---|
| NHANES– KNHANES | Source rows plus labeled target rows; disjoint target query | XGB | 300 trees, depth 6, learning rate .05; source and target total weights balanced | 10 seeds (50–59) per endpoint | Endpoint |
| NHANES/ KNHANES–HRS | All eligible source rows (endpoint-specific); 1/5/10% labeled target rows | XGB | 300 trees, depth 5, learning rate .05; target weight 10 | 10 seeds (40–49) per labeled-target fraction | Endpoint |
| CAMELS-US– CAMELS-GB v2 | 1/3/5 labeled target basins; fixed disjoint query basins | XGB | 300 trees, depth 6, learning rate .05; log-target regression; target weight 10 | 20 seeds (20–39) per labeled-target size | Basin |
| Controlled | 2,048 source rows; ; 2,000 query rows | XGB, HistGB, TabPFN, MLP | Primary conditions use unit weights; XGB/HistGB: 300 depth-6 iterations; TabPFN 7.0.0 defaults; MLP: one 32-unit ReLU layer | 20 realizations 3 fixed families | Realization within family |
| Analysis | Timing / status | Interpretation |
|---|---|---|
| Measurement endpoint extension | Frozen extension | The broader ten-endpoint result is descriptive. |
| Examination width-matched reference | Post-result construction | Enables a new predictive-utility comparison; it does not retroactively make the original ablation a suitable reference comparison. |
| Examination target-support ladder | Direction calibrated earlier; six-endpoint panel selected after the 13-endpoint expansion | Six-endpoint trajectory is treated as replication; the full 13-endpoint trajectory is exploratory. |
| Relation capacity analysis | Frozen extension | Evaluates learner-capacity dependence of relation utility. |
| Fixed-width noisy-proxy diagnostic | Post-result diagnostic; protocol fixed before execution | Tests reconstructibility while preserving input width and variable identity; the association within each noise level is post-result and descriptive. |
| Real-data reconstructibility probes | Frozen extension | The reconstruction-manipulation criterion was not met in the evaluated real-data settings. |
| Endpoint | NHANES item | KNHANES item |
|---|---|---|
| Diabetes history | DIQ010 | ALL__de1_dg |
| Hypertension history | BPQ020 | ALL__di1_dg |
| Current smoking status | SMQ040 | ALL__bs3_1 |
| Kidney disease history | KIQ022 | ALL__dn1_dg |
| Stroke history | MCQ160F | ALL__di3_dg |
| Arthritis history | MCQ160A | ALL__dm1_dg |
| Comparison | Removal | Preservation | Interpretation |
| Measurement altered control | ✓ | Altered condition | |
| Permuted correspondence | ✓ | Content sensitivity only | |
| Survey ablation without added columns | ✓ | Descriptive ablation | |
| Separate-domain reference | ✓ | ✓ | Suitable reference |
| Reversed / random relation control | ✓ | Altered conditions | |
| Relation unconstrained reference | ✓ | ✓ | Suitable reference |
| Relation | Dir. | Supporting evidence | Scope of the evidence |
|---|---|---|---|
| Age SBP | ( Franklin et al., 1997 ; Mitchell et al., 2010 ) | Population trends do not imply the same association under every conditioning set. | |
| SBP DBP | ( Zhu et al., 2022 ; Franklin et al., 1997 ) | The population association is positive, although DBP changes differently with age. | |
| Fasting glucose HbA1c | ( Nathan et al., 2008 ) | The cited relationship uses average rather than fasting glucose. | |
| Total cholesterol triglycerides | ( Friedewald et al., 1972 ; Wolska et al., 2026 ) | Their joint use in LDL estimation is not direct evidence of a general population correlation. | |
| Weight waist circumference | ( Flegal et al., 2009 ; Ross et al., 2020 ) | Much of the cited evidence concerns BMI rather than weight alone. | |
| BUN creatinine | ( Hosten, 1990 ; Morgan et al., 1977 ) | Both reflect renal function, but their association is not uniformly strong across populations. |
| Dataset | Content sensitivity | Predictive utility | Reference–altered |
|---|---|---|---|
| Bank | -.032 [-.044,-.019] | +.089 [+.078,+.101] | -.121 [-.136,-.106] |
| Blood | +.038 [-.020,+.095] | +.114 [+.032,+.197] | -.076 [-.156,+.003] |
| California | +.084 [+.070,+.097] | +.053 [+.039,+.066] | +.031 [+.018,+.042] |
| Car | +.409 [+.370,+.448] | +.300 [+.268,+.333] | +.108 [+.068,+.149] |
| Credit-g | +.066 [+.005,+.129] | -.003 [-.047,+.039] | +.069 [+.015,+.124] |
| Diabetes | +.087 [+.053,+.122] | +.164 [+.102,+.224] | -.077 [-.138,-.016] |
| Panel | Content sensitivity | Predictive utility | Reference–altered | Reference AUROC |
|---|---|---|---|---|
| Focused panel (3 endpoints) | +.292 [+.145, +.431] | +.384 [+.202, +.553] | -.093 [-.228, +.012] | .450 |
| Extended panel (10 endpoints) | +.249 [+.182, +.309] | +.197 [+.110, +.294] | — | .556 [.499, .608] |
| Quantity | Comparison | Mean [95% CI] |
|---|---|---|
| Content sensitivity | Intended – altered correspondence | +.070 [+.029, +.121] |
| Predictive utility | Intended – width-matched reference | +.047 [+.017, +.084] |
| Reference–altered | Width-matched reference – altered correspondence | +.022 [+.002, +.043] |
| Setting | Reversed sensitivity | Random sensitivity | Predictive utility |
|---|---|---|---|
| NHANES–KNHANES | +.183 [+.081, +.298] | -.001 [-.006, +.003] | -.002 |
| NHANES/KNHANES–HRS | +.266 [+.175, +.367] | +.041 [+.016, +.077] | +.002 |
| CAMELS-120 | +.144 [+.127, +.163] | +.002 [+.000, +.003] | .000 [-.001, +.001] |
| Dataset | Primary | Minimum | Maximum | Span |
|---|---|---|---|---|
| Bank | +.089 | +.022 | +.166 | .144 |
| Blood | +.114 | +.093 | +.246 | .153 |
| California | +.053 | -.011 | +.105 | .116 |
| Car | +.300 | +.270 | +.345 | .075 |
| Credit-g | -.003 | -.040 | +.045 | .086 |
| Diabetes | +.164 | +.088 | +.217 | .130 |
| Model | Spearman | Full-mismatch utility | |
|---|---|---|---|
| XGB | 32 | .78–.90 | +.259–+.260 |
| HistGB | 32 | .79–.89 | +.257–+.264 |
| TabPFN | 32 | .75–.85 | +.176–+.249 |
| XGB | 512 | .55–.77 | +.028–+.037 |
| HistGB | 512 | .63–.76 | +.030–+.035 |
| TabPFN | 512 | .57–.80 | +.022–+.035 |
| Predictive utility | 95% CI | |
|---|---|---|
| 0 | +.088 | [+.034,+.153] |
| 4 | +.073 | [+.012,+.144] |
| 32 | +.035 | [+.017,+.054] |
| 512 | +.007 | [+.003,+.011] |
| Decrease, | +.081 | [+.025,+.148] |
| Endpoints | Decrease, | |||
|---|---|---|---|---|
| Six endpoints | +.069 [+.037,+.102] | +.047 [+.018,+.084] | +.018 [+.008,+.028] | +.052 [+.028,+.077] |
| All 13 endpoints | +.039 [+.007,+.069] | +.014 [-.013,+.042] | +.008 [-.001,+.017] | +.031 [+.003,+.056] |
| Setting | Original utility | Rank-aligned utility | Utility increase per reversal |
|---|---|---|---|
| , XGB | +.259–+.260 | +.169–+.275 | +.088 [+.046,+.129] |
| , HistGB | +.257–+.264 | +.167–+.278 | +.088 [+.046,+.129] |
| , XGB | +.028–+.037 | +.008–+.019 | +.004 [+.001,+.008] |
| Family | Depth 6 | Depth 2 | |
|---|---|---|---|
| Additive | 32 | +.0244 | +.0017 |
| Additive | 512 | +.0212 | +.0012 |
| Pairwise | 32 | +.0096 | +.0003 |
| Pairwise | 512 | +.0105 | +.0003 |
| Sparse | 32 | +.0068 | -.0014 |
| Sparse | 512 | +.0067 | -.0015 |
| Model / support | Component | Result |
|---|---|---|
| TransTab / | Shared cross-domain identity | Utility – ; 74–80% of the intended–altered gap is associated with altered-condition degradation |
| TransTab / | Shared cross-domain identity | Utility – ; 97–100% of the intended–altered gap is associated with altered-condition degradation |
| CARTE / | Measurement / correspondence / relation | Positive utility in 0/3, 0/3, and 0/3 outcome families, respectively |
| CARTE / | Measurement / correspondence / relation | Positive utility in 2/3, 2/3, and 0/3 outcome families, respectively |
| Setting | Units | Content sensitivity | Utility / ablation |
|---|---|---|---|
| Survey correspondence / XGB | 3 endpoints | all positive; +.022–+.072 | -.010–+.076 |
| Survey correspondence / HistGB | 3 endpoints | all positive; +.025–+.089 | -.024–+.073 |
| NHANES–KNHANES relation | 6 endpoints | all positive; +.010–+.445 | -.013–+.007 |
| NHANES/KNHANES–HRS relation | 7 endpoints | all positive; +.068–+.531 | -.014–+.019 |
| CAMELS relation | 12 basins | all positive; +.032–+.463 | -.004–+.008 |
| Study | Included | Basis for decision |
|---|---|---|
| LIFT | Yes | Explicit feature-name semantic manipulations |
| TransTab | No | Uses column and categorical descriptions, but no reported ablation alters or removes this semantic knowledge |
| TabLLM | Yes | Explicit name/value semantic manipulations |
| PLATO | Yes | Explicit knowledge-graph ablations |
| CARTE | Yes | Explicit semantic-representation ablations |
| FeatLLM | Yes | Explicit feature-description ablation |
| Coding item | Agreement | Cohen’s |
|---|---|---|
| Content-sensitivity claim | 25/25 (100%) | 1.000 |
| Predictive-utility claim | 24/25 (96%) | .907 |
| Content-sensitivity support | 24/25 (96%) | .815 |
| Removal | 25/25 (100%) | 1.000 |
| Preservation | 15/25 (60%) | .265 |
| Setup | 22/25 (88%) | .784 |
| Content sensitivity | Predictive utility | ||||
|---|---|---|---|---|---|
| Study | ID | Claim | Support | Claim | Support |
| LIFT | C01-A | Y | S | Y | U |
| LIFT | C01-B | Y | S | Y | U |
| LIFT | C01-C | N | – | Y | U |
| LIFT | C01-D | N | – | Y | I |
| TabLLM | C03-A | Y | S | N | – |
| Study | ID | Claim | Claim basis | Source |
|---|---|---|---|---|
| LIFT | C01-A | Y | The gain over shuffled names is attributed to correct feature–value association in prompt format I. | Sec. 4.1 / Table 9, p. 8 |
| LIFT | C01-B | Y | The same interpretation attributes the gain in prompt format II to correct feature–value association. | Sec. 4.1 / Table 9, p. 8 |
| LIFT | C01-C | Y | Correctly incorporating feature names is stated to improve performance except on CMC; this interpretation also covers the no-name comparison. | Sec. 4.1 / Table 9, p. 8 |
| LIFT | C01-D | Y | The same qualified feature-name benefit is stated for the second no-name comparison. | Sec. 4.1 / Table 9, p. 8 |
| PLATO | C04-A | Y | Broader domain knowledge is interpreted as improving prediction beyond a feature-only knowledge graph. | Sec. 4.1 / Table 3, pp. 7, 9 |
| PLATO | C04-B | Y | Feature information in the knowledge graph is interpreted as improving prediction relative to the no-KG configuration. | Sec. 4.1 / Table 3, pp. 7, 9 |
| Component | Version |
|---|---|
| Python (principal analyses) | 3.10.19 |
| NumPy | 2.2.6 |
| pandas | 2.3.3 |
| scikit-learn | 1.7.2 |
| XGBoost | 3.1.3 |
| TabPFN | 7.0.0 |
| Artifact directory | Contents |
|---|---|
| provenance/ | complete record of when each analysis was designed and the associated decisions |
| tabllm/ | complete frozen reference library, reference-construction records, reference-sensitivity results, minimum and maximum utility for each dataset, results for all references with 512 labeled adaptation examples, and utility changes with labeled examples for each reference |
| robustness/measurement/ | absolute AUROC results for the focused measurement analysis |
| robustness/correspondence/ | correspondence results and record of the clean rerun |
| robustness/relation/ | evidence supporting the retained relations, results after removing inputs used to compute relation values, and checks of how well relation values can be recovered |
| robustness/interaction/ | complete interaction results, endpoint robustness checks, label corruption checks, and source-data integrity checks |