DSReg: Provably Recovering Individual World Latents without Reconstruction
Authors: Yujia Zheng, David Klindt, Randall Balestriero, Bernhard Schölkopf
Organizations: University of Illinois Urbana-Champaign · Cold Spring Harbor Laboratory · Brown University · Max Planck Institute for Intelligent Systems · ELLIS Institute Tübingen
Methods that recover individual latent variables of the world, from nonlinear ICA to dictionary learning and causal representation learning, anchor the latents to observations through reconstruction, auxiliary supervision, or distributional asymmetries such as non-Gaussianity. Methods without these anchors, including joint-embedding predictive architectures (JEPAs), identify the latent state only up to a linear transformation, so individual latents remain mixed. We close this gap: individual world latents can be provably recovered with no reconstruction, no decoder, and no labels. The key condition is Structural Diversity: different latents leave distinct dependency footprints on observations, just as no two snowflakes are alike. Building on the linear identifiability that LeJEPA provides, we prove that under Structural Diversity, DSReg (Dependency-Sparsity Regularization) recovers individual world latents up to signed permutation, without reconstruction or a decoder. It applies post hoc to any linearly identified representation, reusing trained checkpoints at no loss over joint training, and establishes the first fully identifiable JEPA that recovers every world latent. Moreover, as a condition on dependency footprints, Structural Diversity is strictly weaker than all structural conditions of prior identifiable latent variable models. Across synthetic regimes, world model probes, learned visual encoders, and external renderers, DSReg preserves dense prediction while improving individual-latent recovery and downstream use with scales.
Figures & tables
Figure 1: DSReg turns a mixed latent state into individual latents. World latents z generate observations x=g(z) . LeJEPA identifies the latent state up to a linear transformation, z^=Az , so each learned variable can remain a mixture of ground-truth latents. DSReg recovers individual latents up to signed permutation.
Figure 2: Dependency footprints. Each column of ∂x/∂z marks which observed variables one latent affects.
Figure 3: DSReg recovers individual latents wherever Structural Diversity holds, and only there. (a) DSReg (solid) is near ceiling at every N , twenty runs. LeJEPA (dashed) stays mixed. (b) On analytic orbits h=Qz , recovery succeeds wherever the condition holds and falls back to the shared pair subspace where it fails.
Figure 4: DSReg improves all six sparse-use probes. DSReg in teal, LeJEPA in gray. Tasks are (1) editing (top) and (2) control, prediction, and monitoring (bottom).
Figure 5: DSReg matches the supervised Procrustes oracle on encoders trained from pixels. (a) Latent-only rotations leave the variables mixed. DSReg closes the gap to the label-fitted oracle (fifteen seeds). (b) With dimz=8 fixed, the match holds within 0.001 as the estimate widens to dimz~=32 (five seeds).
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
MCC
step time (ms)
peak memory (GB)
d
n
mat.
fact.
samp.
mat.
fact.
samp.
mat.
fact.
samp.
128
20 k
0.853
0.853
0.850
7.5
4.9
42.6
0.14
0.11
0.11
256
20 k
0.843
0.844
0.840
7.4
6.8
43.9
0.47
0.31
0.30
512
24.6 k
0.833
0.831
0.829
13.1
34.0
43.8
1.76
1.07
1.02
1024
49.2 k
0.827
0.828
0.823
101.6
231.6
53.5
6.97
4.20
3.26
2048
98.3 k
0.821
0.820
0.809
699.8
1933.2
92.5
27.76
13.73
11.42
Appendix
Table 1 : Scaling across latent dimension. The full estimated procedure, three seeds each. At d≥1024 the sampled columns use a compute-matched budget whose wall-clock stays at or below the materialized one; at smaller d the standard budget already suffices, and sampling holds no advantage there since its per-step cost is nearly d -independent while the materialized cost is what grows. All three modes optimize the same objective, and the materialized and factored columns agree on MCC to within 0.002 everywhere, which is a live check that the factorization is exact. The factored mode trades time for memory, roughly halving peak memory at d=2048 at about 2.8× the step time, since it recomputes the neighbor differences on each pass. Axes without a panel of their own, each measured on its own sweep with its own baseline rather than on the rows above: raising the sample count to 106 leaves recovery and step time flat to within seed noise ( 0.826→0.833 at d=1024 , 49 ms throughout; 0.822 at d=2048 ), and growing the observed dimension to 217 holds the sampled step at 47 ms while MCC improves from 0.806 to 0.884 across the full doubling sweep.
Figure 6 : Scaling the full estimated procedure. (a) Recovery declines gently with d ; open markers are the single-GPU frontier, where anchor count shrinks to fit memory. (b) Recovery against the neighborhood ratio k/d at three scales; dashed lines mark the analytic score of an unrotated mixing, 2lnd/d . (c) Sixteen anchors already match 256 at d=2048 . (d) The sampled step cost is nearly flat in d while the materialized cost grows two orders of magnitude, with the curves crossing between d=512 and d=1024 .
Figure 7 : Footprint regimes across dimension. Recovery follows Structural Diversity at every dimension: diverse, nested, and minimal-difference footprints stay at ceiling while identical footprints sit at the shared-subspace level, far above the decaying baseline of the unrotated representation; the identical pair’s span is still recovered, with CCA at least 0.998 at every dimension of the full regime sweep.
Figure 8 : Recovery is insensitive to the optimizer’s choices. Recovery is insensitive to the sparsity surrogate and to the anchor budget; only the minimal-difference regime rewards more anchors, exactly where the evidence is scarcest. The two rightmost bars replace the relaxation by a hard support count at τ=0.05 , minimized by coordinate search from the same six starts, and by that count used to refine the ℓ1 solution afterwards. Both series are DSReg selections, with color denoting the underlying footprint regime.
Schedule
Latent MCC
LeJEPA (no rotation)
0.640±0.041
Joint
0.906±0.014
Joint, then decoupled fit
0.922±0.005
Decoupled (default)
0.921±0.005
Appendix
Table 2 : Training-schedule ablation on Gaussian 3DShapes. Latent MCC (mean ± std), twenty runs per arm on a shared encoder trajectory; the rotation head never feeds back into the encoder.
Figure 9 : Corrupted Jacobians leave recovery at ceiling. Diverse footprints stay above 0.99 through noise 0.20 , and identical footprints stay at their predicted shared-subspace level. Nonlinear observation rows, so Assumption 2 holds; N=16 , five seeds. LeJEPA applies no rotation, so its curve is flat by construction.
Figure 10 : Sparse access improves while dense content is unchanged. At identical dense recovery ( R2=0.988 for both methods), the rotation transforms individual-latent recovery, few-shot readout from eight labels, and sparse editing, each moving decisively in DSReg’s favor on the MLP benchmark at N=8 .
Figure 11 : Success strictly contains the guarantee. Success fraction of the thresholded criterion over the (τ,ε) grid, five seeds. Solid curve: the guarantee boundary τ+ε=σρ∗(δ∗) of Theorem 3 , with seed-range shading; circled cells satisfy the guarantee for every seed, and all of them succeed. The ℓ1 rotation used throughout the paper reproduces this map on all but four of the 108 cells, again succeeding on every guaranteed cell.
Figure 12 : The excluded linear stratum is thin. MCC against the nonlinearity share γ with the measured no-cancellation margin σ overlaid (dotted, right axis). Where Assumption 2 fails ( γ=0 , σ=0 ) the criterion returns a mixed rotation; recovery crosses the 0.9 level (dashed) at γ∗≈0.1 and reaches the ceiling by γ≈0.35 at every dimension, in both regimes, with exact and estimated Jacobians in close agreement.
Benchmark
N
Dense R2
MCC
Few-shot R2
Sparse error ↓
Visual factors
8
0.918±0.002
0.625→0.958
−0.081→0.882
2.13→1.35
Object scenes
8
0.924±0.002
0.591→0.961
−0.143→0.899
2.29→1.41
Appendix
Table 3 : Split-audited learned visual encoders. The analytic signature survives learning: dense recovery is unchanged while individual-latent recovery, few-shot readout, and sparse use improve, with every score computed on a held-out split (the rotation is fit on a separate split of estimated latents and fixed pooled-pixel features). Arrows are ten-run means for LeJEPA → DSReg; sparse error is lower better.
Figure 13 : The analytic signature survives pixel training. The gain concentrates where individual latents matter: component-wise recovery and single-latent readout improve sharply, and sparse use follows. A negative few-shot value underperforms the held-out mean; the sparse-use score is normalized so that LeJEPA sits at one. On the pixel probes, sparse control error drops from 0.63 to 0.52 , five-step sparse rollout R2 rises from 0.74 to 0.81 , and transition-surprise AUROC from 0.995 to 0.999 (LeJEPA / DSReg, ten runs).
Figure 14 : Training imbalance does not break the rotation. Long-tail and correlated training distributions leave the recovery gain intact, with the dense span fully preserved throughout.
Method
Latent MCC
DCI
MIG
SAP
Info. R2
Dense R2
LeJEPA
0.640±0.041
0.298±0.044
0.069±0.026
0.212±0.074
0.854±0.003
0.861±0.002
β -VAE
0.272±0.025
0.318±0.086
0.080±0.038
0.102±0.030
0.738±0.076
0.191±0.010
β -TCVAE
0.277±0.043
0.345±0.117
0.102±0.067
0.109±0.052
0.751±0.087
0.199±0.027
DSReg
0.921±0.005
0.901±0.017
0.419±0.013
0.842±0.016
0.868±0.003
0.861±0.002
Appendix
Table 4 : DSReg improves latent MCC on every run of Gaussian 3DShapes while matching LeJEPA’s dense recovery. Scores are latent MCC, DCI, MIG, SAP, and informativeness (test R2 of a boosted readout), twenty runs, on an external renderer. The LeJEPA row is the same frozen representation DSReg rotates, which is why the two share dense R2 ; the β -VAE/ β -TCVAE rows are external references with the matched trunk, trained on i.i.d. images at an untuned β=4 ; their low dense R2 reflects a linear readout only, with nonlinear informativeness staying high, so the gap measures axis alignment rather than lost information.
Figure 15 : Per-factor recovery on Gaussian 3DShapes. DSReg improves every factor and alone recovers scale, shape, and orientation, where the VAE baselines collapse almost entirely.
Figure 16 : The rotation turns entangled responses into one selective response per factor. One run of Gaussian 3DShapes. Each row sweeps one factor with all others fixed (rendered frames on the left); the right panels show every learned latent’s standardized response in LeJEPA and in DSReg coordinates, on a shared scale per row. The flat gray curves carry the no-mixing claim across all six factors.
Figure 17 : Matched correlation heatmaps. For the run of Figure 16 , columns are permuted by each method’s own best assignment. The same encoder yields a smeared matrix in the unrotated coordinates and a near-diagonal one after rotation, in agreement with the per-factor responses of the figure above.
Figure 18 : DSReg beats both VAE baselines under a partially attained premise. Quarter-orientation dSprites, twenty runs, dense R2≈0.53 as the rotation-invariant ceiling; each VAE at its best β . One-sided Mann–Whitney tests are against β -VAE and β -TCVAE at their own best β ; higher MCC is better.
Method
Latent MCC
R2(h→z)
Dependency sparsity ↓
LeJEPA
0.65±0.03
1
0.32±0.02
Random orbit
0.64±0.03
1
–
PCA
0.62±0.04
1
–
Varimax
0.64±0.03
1
–
FastICA
0.64±0.05
1
–
DSReg
0.90±0.05
1
0.18±0.02
Appendix
Table 5 : Latent-only rotation baselines. The observation-side signal rather than the span is what does the work: all methods start from the same linearly identifiable representation, but PCA, Varimax, and FastICA see only estimated-latent samples and cannot break the rotational symmetry, while the dependency criterion ∂x/∂(Rh) can, because it also sees how the observed variables respond to the applied rotation.
Task
Retrieval cost LeJEPA / DSReg
Sparse err. LeJEPA / DSReg
TwoRoom
1.31/1.31
0.81/0.58
PushT
1.45/1.45
1.45/1.10
FetchSlide
1.43/1.43
1.24/0.85
Appendix
Table 6 : Dense versus sparse-use sanity check. Everything the rotation gains, it gains in sparse access: dense readout and ordinary retrieval are unchanged, with dense R2 at 1/1 (LeJEPA / DSReg) for every task, while sparse-control error drops whenever only a few estimated latents may be used.
Criterion
Gaussian 3DShapes
Quarter-orientation dSprites
Column-support sparse ICA
0.76±0.06
0.57±0.05
DSReg
0.92±0.01
0.60±0.08
Appendix
Table 7 : Sparse ICA on the same dependency signal recovers less. Latent MCC, mean ± std over ten fresh runs per renderer benchmark; both criteria consume the same estimated anchor Jacobians, orthogonal parameterization, and restarts. Both rows share every setting except the criterion.