Complementary Feature Domains (CFD) theory characterizes predictive value as a context-indexed contribution system induced jointly by representations and their realization family. We show that Shannon-information preservation does not imply preservation of this contribution system: an invertible representation transformation can leave target information unchanged while altering predictive contribution under a restricted decision family. We formalize the resulting transition through a CFD contribution defect that measures how contextual contributions change under controlled recoding. For bounded Lipschitz utility, we show that each coalition utility shift is bounded by the behavioral distance between the attainable action sets before and after recoding; consequently, every contextual contribution defect is bounded by the sum of the corresponding coalition incompatibilities. Exact behavioral closure yields invariance, while increasingly accurate compensation yields restoration. A controlled ECG experiment illustrates the mechanism: a nonlinear bijective recoding preserves the information in a frozen time-frequency representation but changes accuracy under a fixed affine learner; applying the exact inverse restores all tested coalition accuracies. The result separates information preservation from realization-dependent contribution and provides a quantitative transition law for multi-representation prediction.
Figures & tables
Framework
Primary object
Question answered
MI/CMI [ 2 ]
Target dependence
How much target information is present?
PID [ 4 ]
Information decomposition
How is target information distributed among sources?
V -information [ 9 ]
Family-relative usable information
How much information can a predictor family exploit?
Blackwell [ 10 ]
Statistical experiments
Which information structure is at least as useful for decisions?
CFD
Contribution transition
Change of contextual roles under a representation–realization intervention
TABLE I: Analytical objects and the distinct CFD transition question.
Realization family
XOR
Recoded
Gap
Affine (exact optimum)
.75
1.00
.25
Quadratic logistic
1.00
1.00
0
RBF-kernel SVM
1.00
1.00
0
Decision tree, depth 2
1.00
1.00
0
TABLE II: Capacity stress test. Both encodings contain exactly one bit of target information.
Target
VB(12)
VB(123)
Δ3,B(12)
Δ3,A(12)
Parity
.500
1.000
.500
.250
Majority
.750
1.000
.250
.250
OR
.875
1.000
.125
.125
AND
.875
1.000
.125
.125
TABLE III: Exact full-context contribution under Bayes (B) and affine (A) realization.
Coalition
Raw
Recoded
Restored
Recoding gap
TF
.7372
.6277
.7372
−.1095
T+TF
.6569
.5839
.6569
−.0730
F+TF
.7883
.5912
.7883
−.1971
T+F+TF
.7372
.6058
.7372
−.1314
TABLE IV: Information-preserving recoding and inverse restoration. Intervals are paired 95% integrity-group bootstrap intervals for recoded minus raw accuracy.
Prediction models trained under the source distribution do not generalize well to a different target distribution. A valid inference about an unseen data distribution must be anchored by the invariance of certain causal mechanisms that generate the source and target data, however, these structural invariances are non-identifiable from the source data alone. Under mild causal assumptions about the data, we show that the optimal prediction in the target is in fact partially identifiable by the source distribution. The result rests on a simple observation: In any domain, the optimal prediction can be factorized into what we call a pair of abduction and deduction maps, where the abduction map makes inference about some unobserved variables (possibly confounders) from the observed variables and the deduction map predicts the label using both the observed and inferred quantities. Access to large source data pins down the optimal prediction, thus constrains the valid abduction-deduction ensembles that produce it -- a non-identifiability that we call the abduction-deduction entanglement. To leverage this, we parameterize the constrained family using what we call a representation transplant, that is a specific linear transformation in the representation space that manipulates the abduction content of the representation while retaining the deduction component. Invariance of the causal mechanism generating the label implies existence of an invariant deduction map between source and target. Thus, we can search the space of plausible target distributions via a parametric transplant. We use this scheme in a learner-adversary game that, under an idealistic optimization, provably terminates with the learner having the minimax-optimal target prediction. Evaluations verify the theory, showing that the method is competitive in DG benchmarks.
Kasra Jalaldoust, Elias Bareinboum
Department of Computer Science Columbia University New York, NY 10027, USA
We report a systematic failure mode in predictive representation learning. Across 2695 neural network configurations trained to predict linear-Gaussian dynamics, the optimal encoder tracks the environment rather than the system it is meant to model. The mean causal fidelity -- the fraction of encoder sensitivity allocated to system degrees of freedom -- is 0.49, and only 2.5% of configurations exceed 0.70. The failure intensifies with dimension: at N=100, the optimal encoder becomes causally blind (fidelity ~10^{-8}) while achieving 92% lower prediction error than the causal representation. We prove this is not an optimization artifact but a structural property of the predictive objective: when environment modes are slower or less noisy than system modes, every minimizer of the population risk encodes the former. The set of dynamics exhibiting this predictive-causal gap is open and of positive measure in parameter space. In a nonlinear Duffing-GRU sweep, unconstrained predictors learn environment-dominant representations in 55% of tasks (95% CI 41--68%) versus 24% under operational grounding (p=2.3e-3); the median out-of-distribution MSE inflation under environment shift is 1.82x versus 1.00x. Operational grounding -- restricting the loss to system observables -- partially suppresses the gap, but causal fidelity is never recovered without an explicit system-environment boundary. The results identify the predictive-causal gap as a structural limit of learning, with implications for self-supervised representation learning, world models, and the scaling paradigm.
Kejun Liu
State Key Laboratory of Bioinspired Interface Material Science, Institute of Nano & Functional Materials, Soochow University, Suzhou 215123, China
Estimating counterfactual outcomes over time from longitudinal observational data is central to clinical decision support. Existing methods rely on domain confusion -- adversarial training that renders representations invariant to treatment assignment -- yet this invariance creates a mutual information conflict: it suppresses treatment-correlated covariate signals necessary for accurate outcome prediction. We formalise this tension via a Jensen-Shannon divergence bound on counterfactual prediction error and develop two complementary models. CSSD (Causal State-Space model with Direct decoder) adapts selective State Space Models with a parallel multi-step decoder that eliminates accumulated rollout error by producing all prediction horizons simultaneously in a single forward pass. CSSPD (Causal State-Space model with Predictive regularisation and Direct decoder) augments CSSD with Contrastive Predictive Coding and Local Information Maximisation to reinforce temporal predictability in the balancing representation and recover local covariate information destroyed by domain confusion. On MIMIC-III, CSSPD achieves lower counterfactual RMSE than the Causal Transformer at every horizon tau >= 2 at O(T) encoder cost, with gains from 0.02 (2-step) to 0.07 (6-step). On Cancer Simulation across confounding strengths gamma in {0,1,2,3,4}, CSSPD outperforms CT at gamma <= 3 (margins 25.9%--37.0%), and CSSD achieves the lowest overall average RMSE (12.7% reduction over CT), confirming the MI conflict analysis. To our knowledge, this is the first work to formalise the balancing-prediction MI conflict and propose a structured resolution through complementary predictive and information-theoretic training objectives.