Embedding-Bias in Conditional Independence Testing
Organizations: Department of Mathematical Sciences University of Copenhagen Copenhagen, Denmark
Abstract
To test conditional independence of and given a text or an image , one conditions on an embedding in place of . The embedded test is valid if is independent of or of given , which cannot be confirmed from data, and when this fails, the rejection probability under the null hypothesis can tend to one. We study this failure, and show that focusing on a specific form of dependence relaxes what the embedding must retain. For a residual correlation test inspired by the Generalised Covariance Measure, validity only requires that the parts of and missed by and are uncorrelated. Otherwise, we treat the discarded information as an omitted variable. Under the null hypothesis, the bias equals the absolute correlation of the missed parts times the geometric mean of two partial values. This identity yields a robust test valid under a declared tolerance for the geometric mean, which, like a sensitivity parameter, is not identified from the data. On synthetic data and text embeddings, the robust test holds its level approximately. On text generated by a language model, under an exact null hypothesis, every embedding, even the generator's own states, biases the embedded test.
Figures & tables
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Explore similar work
Meta-Dependence in Conditional Independence Testing
meta-dependence'' between conditional independence properties using the following geometric intuition: satisfying each conditional independence property constrains the space of possible joint distributions to a manifold. The meta-dependence'' of multiple conditional independences in a probability distribution is informed by its position relative to these manifolds. We provide a simple-to-compute measure of this meta-dependence using moment projections, with a closed-form expression for multivariate Gaussian distributions, and consolidate our findings empirically using both synthetic and real-world data. Our measure of meta-dependence does not rely on graphical properties of the distribution and can be computed directly from summary statistics such as a covariance matrix, allowing for various applications. We demonstrate one use case of meta-dependence, using a simple redundancy metric to tune significance thresholds and improve causal discovery.