cs.LGJun 29, 2026

Exploring Differences Between Tabular Enterprise Data and Public Benchmarks

Authors: Myung Jun KimMaximilian SchambachFrank EssenbergerAndre SresJohannes Höhne

Abstract

Tabular data dominate the landscape of data science, increasingly attracting innovative machine learning models and tailored benchmarks. Yet, little is known for enterprise data, where tables constitute the backbone of business operations. To broaden the benchmarking landscape for business applications, this work aims to actualize the characteristics of enterprise data by providing an analysis of data statistics and performance measurements of tabular models such as TabPFN, TabICL and ConTextTab. Through our analysis, we find enterprise data markedly differ from tabular benchmarks and we demonstrate that a tabular model that performs well on typical tabular benchmarks may perform poorly on real world enterprise data -- and vice versa. This lack of generalization underlines the need for additional benchmarks with enterprise-grade characteristics.

Explore similar work

CardsList
  1. Towards Universal Tabular Embeddings: A Benchmark Across Data Tasks

    Apr 23, 2026Liane Vogel, Kavitha Srinivas, Niharika D'Souza +3Tabular Learning