CI-JEPA: A Counterfactual Analysis of Latent Representations in Joint-Embedding Predictive Architectures for Self-Supervised Learning
Organizations: Department of Information and Communication Technology, Pandit Deendayal Energy University, Gandhinagar, Gujarat, India · Department of Electrical and Electronics Engineering, Birla Institute of Technology & Science, Pilani, Rajasthan, India
Abstract
Self-supervised visual representation learning learns useful features without manual annotations during representation training. The image-based joint-embedding predictive architecture (I-JEPA) predicts latent representations of masked image regions, but its objective does not explicitly model responses to specified visual interventions. We introduce CI-JEPA, a counterfactual intervention-aware extension that learns to predict the representation change between an original image and a modified counterpart. We assess representation robustness through selective sensitivity: stronger responses to task-relevant semantic changes than to nuisance changes. Experiments on Flowers102 use flower-center occlusion as a candidate semantic intervention and background blur and tint as candidate nuisance interventions. With frozen-encoder linear probing, CI-JEPA achieves a best validation accuracy of 78.14%, compared with 77.55% for both the pretrained ViT-B/16 and the I-JEPA baseline, a gain of 0.59 percentage points. The reported mean representation changes are 4.42 for center occlusion, 3.48 for background tint, and 2.83 for background blur. This ordering is consistent with relative semantic selectivity for the evaluated interventions, rather than complete nuisance invariance. The accuracy comparison is complementary and does not establish improved robustness over the baselines. These controlled image modifications provide a framework for studying intervention-induced changes in JEPA representations; they do not establish causal feature discovery or robustness to all visual changes.
Figures & tables
| Model | Best validation accuracy (%) |
|---|---|
| Pretrained ViT-B/16 | 77.55 |
| I-JEPA | 77.55 |
| I-JEPA + CF | 78.14 |
| Intervention | MAE | RMS | Mean | Std. |
|---|---|---|---|---|
| change | change | |||
| Background blur | 0.08 | 0.11 | 2.83 | 0.74 |
| Background tint | 0.10 | 0.13 | 3.48 | 0.98 |
| Center occlusion | 0.12 | 0.16 | 4.42 | 0.37 |