cs.LG · 2606.12277 Copy arXiv ID · Jun 10, 2026 Save Finding Multiple Interpretations in Datasets Authors: Matthew Chak , Paul Anderson
Organizations: Department of Computer Science California Polytechnic State University San Luis Obispo, CA, US
Abstract In this paper, we propose an approach to finding sets of similar-performing models (in terms of loss/accuracy measurements) with highly different context-aware characteristics. Through experiments on the METABRIC dataset, we show that the proposed method finds multiple models with highly different gene expressions than those found by the control methodology without performance penalties. We argue that the proposed methodology is important whenever one aims to analyze any global characteristic of a model to extract insight into the underlying phenomenon being studied.
Explore similar work May 27, 2026 · Markus Herre, Andrej Tschalzev, Sascha Marton +1 Tabular Foundation Models
May 28, 2026 · Andrada Gobeaja, Ionut Hodoroaga, Elena Burceanu +1 Membership Inference Cross-Dataset Benchmark
May 14, 2026 · ML Nissen Gonzalez, Melwina Albuquerque, Laurence Wroe +3 Mechanistic Interpretability Semantic Similarity Metrics
May 27, 2026 · cs.LG J/K move · Enter open · S save
Markus Herre, Andrej Tschalzev, Sascha Marton, Christian Bartelt
With the rise of tabular foundation models alongside traditional models still performing well on many tasks, choosing the right model for a tabular dataset remains difficult. We investigate whether dataset meta-features can explain performance gaps between model families on tabular prediction tasks. Using the TabArena benchmark results, we analyze dataset-level performance gaps and relate them to model-agnostic meta-features. After strict statistical tests with false discovery control, we find that (1) for neural network vs. tree gaps, no meta-feature survives false discovery control, (2) for non-foundation vs. foundation model gaps, one association is robust but does not generalize when tested in leave-one-dataset-out prediction, and (3) for TabICLv2 vs. TabPFN-2.6, one robust association also improves held-out prediction. Furthermore, we conduct a leave-one-dataset-out analysis and find that meta-feature predictors fail to improve meaningfully over a simple baseline. Overall, our results show the heterogeneity of tabular datasets and that global meta-feature approaches are not robust enough to offer explanations on the 51 TabArena datasets.