Transactional basket data can reveal associations among items, but observed co-occurrence conflates item-specific relations with basket-size structure and unmodeled higher-order dependence. We introduce Cardinality-Stratified Interaction Decomposition (CSID), an interpretable framework that decomposes log-odds contrasts stratified by the number of remaining items into item-set-specific and cardinality-common components, without fitting a global joint distribution. CSID uses an information-weighted, gauge-constrained ridge projection to estimate pair and triple components and to diagnose higher-order contributions to pairwise structure. CSID is designed primarily for interpretable decomposition of association structure rather than for full-distribution prediction. In a simulation with zero pair effects, increasingly strong small-basket cardinality potentials drive ordinary Ising couplings spuriously negative, whereas CSID pair estimates remain centered near zero. Detection power rises with the magnitude of planted triple effects, and local deprojection reduces pair-coefficient RMSE from 0.244 to 0.073. Across three grocery datasets, high-information triple components are reproducible over time. In the matched cross-period partial-transfer evaluation, transferred CSID triple components show closer agreement with later-period stratified contrasts than the nodewise-symmetrized cardinality-aware higher-order pseudolikelihood comparator, with gains in weighted Lin's concordance correlation of 0.038--0.122. These results support CSID as an exploratory and interpretable decomposition framework for pairwise and higher-order association structure in transactional data.
Figures & tables
s
Mean size
Ordinary Ising
CSID-2
Mean estimate
Neg. frac.
RMSE
Mean estimate
Neg. frac.
RMSE
0.0
4.27
-0.001
0.514
0.042
0.000
0.502
0.042
0.25
3.36
-0.056
0.887
0.073
-0.000
0.492
0.047
0.5
2.72
-0.137
0.997
0.146
-0.001
0.493
0.052
0.75
2.28
-0.235
1.000
0.242
-0.002
0.496
0.062
1.0
1.97
-0.361
1.000
0.367
-0.003
0.511
0.080
Table 1: Ordinary Ising and CSID-2 estimates under a cardinality-only data-generating process. Mean estimate is the Monte Carlo mean over 20 replicates of the within-replicate average across all (210)=45 fitted pair coefficients. Neg. frac. and RMSE are averaged analogously from their within-replicate pair summaries; all true pair coefficients are zero.
Figure 1: Mean pairwise estimate θij (left) and RMSE to the true zero vector (right) against mean basket size. Mean basket size increases from left to right; stronger cardinality penalties therefore appear farther left, and strength levels s are listed in table 1 . Lines compare ordinary pairwise Ising and CSID-2, with shaded bands showing 95% Monte Carlo intervals over 20 replicates. As baskets become smaller, ordinary Ising couplings shift downward, whereas CSID-2 remains near zero with substantially lower RMSE.
∣θ∣
Power (95% MC interval)
Non-planted rate
AUPRC
AUROC
Sign recovery
0.1
0.075 [0.035, 0.115]
0.032
0.133
0.526
0.719
0.2
0.194 [0.142, 0.246]
0.034
0.259
0.662
0.850
0.3
0.338 [0.235, 0.440]
0.042
0.393
0.766
0.944
0.5
0.700 [0.589, 0.811]
0.038
0.745
0.924
0.988
Table 2: Detection of randomized sparse third-order signals. Non-planted rate is the fraction of non-planted triples with ∣zT∣≥2 . Monte Carlo intervals are two-sided 95% t intervals over the 20 replicate-level detection rates.
Figure 2: Planted-triple detection rate (power) and non-planted exceedance rate against planted effect size ∣θ∣ . The horizontal axis shows absolute planted coefficients; the vertical axis shows the fraction of fitted triples flagged. Shaded bands are 95% Monte Carlo intervals over 20 replicates ( table 2 ). Power increases with ∣θ∣ , whereas the non-planted exceedance rate remains near 3–4%.
Estimator
Gauge-aligned RMSE
Pearson r
CSID-2
0.244
0.693
Reweighted CSID-2
0.073
0.985
Table 3: Recovery of pair coefficients in one controlled 60,000-basket realization with third-order saturation. RMSE is computed against the generating pair coefficients after information-weighted gauge alignment; Pearson r is computed across all 15 item pairs.
Dataset
Baskets
H1 / H2
Selected categories
n111 min.
Complete Journey
72,936
35,087 / 37,849
20 predefined product_category values
0
Ta-Feng
90,258
44,564 / 45,694
Top 20 two-digit PRODUCT_SUBCLASS codes
10
Instacart
373,133
179,784 / 193,349
Top 20 aisles (30,000-user subsample)
30
Table 4: Datasets and temporal splits after category restriction. The column n111 min. gives the minimum total triple co-occurrence count ∑rn111,r required to retain a CSID-3 candidate in H1 ( eq. 26 ); pairs were retained without an analogous cutoff.
Dataset
Pearson (LOCO range)
Sign agreement
Negative top-20 / expectation
Complete Journey
0.572 [0.539, 0.616]
0.677
6 / 1.40
Ta-Feng
0.739 [0.549, 0.768]
0.729
11 / 1.59
Instacart
0.723 [0.664, 0.742]
0.747
9 / 1.40
Table 5: H1–H2 reproducibility of high-information CSID-3 effects. The LOCO range is the minimum and maximum Pearson correlation over the 20 omitted-category evaluations; the models are not refitted. The top information quartile is reselected after each category omission. The reported expectation is k2/N , the expected overlap of two independently selected sets of k=min(20,N) triples among the N evaluated triples.
Dataset
Triples
Layers
Baseline
CA-HOPL
CSID-3
Gap
LOCO range
Complete Journey
285
4,401
0.356
0.403
0.442
0.038
[0.025, 0.049]
Ta-Feng
251
2,472
0.107
0.358
0.480
0.122
[0.083, 0.141]
Instacart
285
3,812
0.390
0.549
0.597
0.049
[0.039, 0.054]
Table 6: Continuous agreement with H2 third-order contrasts for the top H1-information quartile. Baseline, CA-HOPL, and CSID-3 report information-weighted CCC on identical triples and layers; Gap is CSID-3 minus CA-HOPL. The LOCO range is the minimum and maximum gap over the 20 omitted-category evaluations; the models are not refitted.
Weighted correlation ρω
Bias correction Cb,ω
Dataset
Baseline
CA-HOPL
CSID-3
Baseline
CA-HOPL
CSID-3
Complete Journey
0.499
0.524
0.520
0.71
0.77
0.85
Ta-Feng
0.225
0.454
0.512
0.48
0.79
0.94
Instacart
0.495
0.619
0.637
0.79
0.89
0.94
Table 7: Decomposition of weighted CCC for the common-effect baseline, CA-HOPL, and CSID-3 in the top H1-information quartile. Following ρc,ω=ρωCb,ω , ρω is the weighted Pearson correlation between H2 contrasts and each reconstruction, and Cb,ω measures agreement in mean and scale.
Figure 3: Information-weighted Lin’s concordance correlation coefficient (CCC) between H2 stratified third-order contrasts and the common-effect baseline, CA-HOPL, and CSID-3 transfer reconstructions, plotted against the fraction of H2-evaluable triples retained by H1 information. Each panel shows one dataset; all methods use identical triples, layers, H2 common-effect reconstruction, and information weights. No H2 effect-size filter is applied. CSID-3 CCC exceeds CA-HOPL at the primary 25% fraction in every dataset.
Dataset
H1 common
Full CSID-3
Gain
Partial CSID-3
Complete Journey
0.032
0.158
0.127
0.442
Ta-Feng
-0.115
0.268
0.383
0.480
Instacart
0.127
0.418
0.291
0.597
Table 8: Information-weighted CCC for H1-only reconstruction of H2 contrasts in the top H1-information quartile. H1 common uses β3,rH1 alone, Full CSID-3 uses eq. 57 , Gain is their difference, and Partial CSID-3 repeats the primary reconstruction that estimates the common component from H2.
Figure 4: Conditional-odds contrasts Λijℓ=logORij∣xℓ=1−logORij∣xℓ=0 for eight H1-ranked, H2-evaluable triples in each dataset. Each panel shows one dataset; rows list the focal pair and moderator and are ordered by the observed H2 contrast. Circles with horizontal intervals give observed H2 values with approximate 95% cell-count delta-method reference intervals; these intervals do not account for customer clustering or selection of the displayed triples. Triangles give the common-effect baseline and squares give the CSID-3 transfer. A vertical dashed line marks Λ=0 .
Dataset
Pair
CSID-2
Reweighted CSID-2
Δθ(3)
Complete Journey
Milk – Cold cereal
0.711
1.167
0.456
Complete Journey
Soup – Crackers
0.577
1.008
0.431
Ta-Feng
51 – 54
1.067
2.626
1.559
Ta-Feng
73 – 47
-0.288
-1.029
-0.741
Instacart
Fresh herbs – Cereal
-0.639
-1.168
-0.529
Instacart
Packaged cheese – Lunch meat
0.464
0.973
0.510
Table 9: Illustrative pairwise projections showing large positive and negative third-order corrections. The rows were selected for interpretation rather than by a prespecified inferential rule.
Dataset
Correlation with H2
Correlation with full-period
Complete Journey
0.688
0.906
Ta-Feng
0.695
0.904
Instacart
0.814
0.945
Table 10: Cross-period reproducibility of third-order pairwise corrections. Pearson correlations are computed across all 190 focal pairs between corrections obtained by applying H1 triple coefficients to H2 pair cells and corrections based on H2 or full-period fits. All corrections use the same H1-fixed gauge convention, so the table compares relative correction patterns without changing the reference gauge.
Context aisles
Pair ρ
Triple ρ (top 25%)
Deprojection ρ
ΔCCC
20
0.988
0.950
0.893
0.286
40
0.997
0.992
0.984
0.163
80
0.999
0.998
0.996
0.165
All (134)
1.000
1.000
1.000
0.172
Table 11: Instacart robustness to expansion of the surrounding aisle universe. Each row fixes the focal top-20 aisles and expands only the rest-cardinality context. Pair ρ compares all 190 focal θij to the all-aisle reference; triple ρ (top 25%) compares the highest-information quartile of focal θT ; deprojection ρ compares third-order pairwise corrections Δθ(3) . All correlations use the all-aisle analysis as reference. ΔCCC is the CSID-3 minus baseline improvement in information-weighted Lin’s CCC at ∣zTH1∣≥2 .
Figure 5: Coefficient stability under Instacart item-universe expansion. Left: all 190 focal pair coefficients with 20 context aisles ( θij(20) ) versus the all-aisle reference ( θij(all) ). Right: the top information quartile of focal third-order coefficients ( θT(20) versus θT(all) ). Dashed lines mark equality; annotations give Pearson correlation.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Quantity
Default
ε range
λ range
r≥1
Complete Journey
Temporal ρ
0.572
[0.571, 0.572]
[0.571, 0.572]
0.647
CSID–CA-HOPL CCC
0.038
[0.038, 0.039]
[0.037, 0.038]
0.049
Correction ρ
0.688
[0.683, 0.699]
[0.688, 0.689]
0.777
Ta-Feng
Temporal ρ
0.739
[0.736, 0.745]
[0.738, 0.745]
0.779
CSID–CA-HOPL CCC
0.122
[0.121, 0.122]
[0.114, 0.123]
0.123
Correction ρ
0.695
[0.681, 0.721]
[0.687, 0.746]
0.763
Appendix
Table 12: Focused one-factor sensitivity analysis. Temporal ρ is the H1–H2 Pearson correlation in the top information quartile; CSID–CA-HOPL CCC is the partial-transfer gap; Correction ρ compares H1-based and H2-based pair-correction patterns. The r≥1 analysis restores baskets containing one selected category before constructing the additional layer.
Aisles
Mode
Generated
Retained
Time (s)
Storage (MB)
20
Exhaustive
1,140
1,140
6.9
0.6
20
Sparse
1,140
1,140
6.9
0.6
40
Exhaustive
9,880
9,880
16.5
6.8
40
Sparse
9,880
9,880
16.5
6.8
80
Sparse
82,160
79,193
37.6
83.0
134
Sparse
250,000
156,735
49.9
178.2
Appendix
Table 13: Single-process candidate-scaling audit on 373,133 Instacart baskets, conducted on a laptop computer with an arm64 processor. Time is the sum of basket scanning, candidate construction, exact counting, contrast construction, and CSID-2/3 fitting; storage is the combined size of the retained stratified count and contrast arrays, not peak process memory. Sparse generation used frequent-pair triangles and bounded neighborhoods with a 250,000-candidate cap.
Signed pairwise interaction scores fundamentally conflate uniqueness (U), redundancy (R), and synergy (S). We prove this on a minimal 3-way XOR structural causal model: faithful indices such as Shapley-Taylor return zero per pair, whereas projective indices such as Shapley Interaction spread the third-order effect into pair scalars that conflate the three mechanisms. We introduce Stochastic Hi-Fi, a post-hoc, retraining-free predictability decomposition that estimates per-feature U/R/S profiles by interventional masked inference. The estimator provides exact interventional semantics, finite-sample Monte Carlo bounds, strict variance reduction from coupled diamond sampling, and uniform finite-vocabulary convergence. Across tabular SCMs, Stochastic Hi-Fi recovers structure missed by scalar baselines (up to 411x larger interaction-magnitude recovery ratios). It also separates redundant and synergistic heads in the GPT-2 IOI circuit. On NIH ChestX-ray14, Stochastic Hi-Fi matches GradCAM on Pointing Game and improves substantially on Deletion AUC.
Causal discovery seeks to uncover the causal dependencies among variables. For this purpose, we propose an algorithm called Tensor-based Second-order Causal Discovery (TSCD). Its input is a tensor obtained from the covariance matrices of observational and interventional data. Assuming the causal dependencies follow a linear structural equation model on a directed acyclic graph (DAG), TSCD outputs the DAG and the functions on its edges, requiring only that the noise variables are uncorrelated. We also implement a version of the approach for nonlinear models. Our focus on second-order statistics (via the covariance matrices) is motivated by their statistical and computational efficiency relative to higher-order moments, their identifiability relative to first-order statistics, and that they work regardless of whether the variables are Gaussian. We show that TSCD has identifiable causal order and parameters from a number of interventions that is logarithmic in the number of variables. Experiments show that TSCD is robust to noise, competitive with existing methods, and scales to hundreds of variables.
Inherently interpretable classifiers for tabular data typically rely on sparse features, rules, or patterns that users can inspect directly. The marginal feature-screening step common to these methods can discard variables whose predictive value emerges only through joint configurations with other variables. We present Interaction Aware Interpretable Machine Learning (IAIML), a framework that addresses this limitation through three coordinated mechanisms: adaptive per-feature discretization, finite-grid pairwise interaction scoring, and a partitioned explanation budget. Detected interactions are routed through one of two strategies: relaxing the screening filter so that interaction-supported variables enter the pattern search, or constructing explicit pair terms for a sparse downstream classifier. On a 40-dataset panel comprising 24 real-world tabular benchmarks and 16 synthetic interaction stress tests, evaluated under nested cross-validation, IAIML achieves mean AUC within 1.4 points of tuned gradient-boosted ensembles while requiring roughly 14--28 times fewer fitted explanation components. On datasets with strong pairwise interaction structure and low marginal signal, IAIML outperforms all baselines. Among compact interpretable methods, IAIML is comparable to RuleFit in AUC and component count and is less expensive to tune. EBM obtains a small but significant AUC advantage across the full panel, with a substantially larger lookup-table footprint. Performance degrades on datasets requiring higher-order interactions beyond the pairwise scope. Component-isolated ablations confirm that adaptive discretization and interaction-aware admission each contribute incrementally. These results support IAIML as a compact, interaction-aware framework appropriate for settings where bounded explanation size and controlled treatment of feature interactions are design requirements.
Srikumar Krishnamoorthy
Information Systems Area, Indian Institute of Management Ahmedabad, India