Isotropic Yet Undecodable: The Sequential Content-Sufficiency Gap in Latent-Predictive Text Representations
Authors: K. P. Santoso, N. Z. Fadil, F. P. Harsanti, R. V. H. Ginardi, G. N. Iyer
Organizations: School of Computing, National University of Singapore · Department of Information Technology, Institut Teknologi Sepuluh Nopember · Avalon AI · Faculty of Computer Science, Universitas Indonesia
We study sequential content sufficiency by investigating whether a representation retains the ordered target information available in its input. An information-theoretic decomposition separates input ambiguity, representation loss, and readout mismatch. We construct recoverable views where perfect agreement and joint isotropic Gaussianity coexist with zero target information, and establish limits imposed by deterministic canonical anchors. Token log-loss provides a one-sided information-loss bound; a fixed-penalty ridge analysis shows why rank alone cannot determine prediction risk. These results motivate CANOPE, a nonautoregressive framework with ordered latent canvases, canonical-token supervision, and geometric regularization. On 40,000 validation sequences, latent-agreement (PL0) and token-grounded (PL2) have nearly identical pooled ranks but reach 13.5% and 98.8% positional Recall@1, respectively, under strong natural corruption when the correct target length is provided. On 3,930 LJSpeech validation utterances, frozen PL2 with a trained MatchaTTS readout yields 21.54% word error rate (WER) on corrupted text, versus 99.22% for frozen PL0, while end-to-end MatchaTTS reaches 10.93%. These results show that geometric regularity alone does not guarantee recoverable sequential content or effective downstream access in the text settings studied here.
Figures & tables
Figure 1: Latent agreement and canonical-token grounding. The recoverable CL channel isolates access to ordered content; downstream readouts test reconstruction and speech. Both objectives use the same architecture.
Figure 2: CL relative-position retrieval with oracle or input lengths (40,000 targets; level 0: clean). CL-GT is strongly length-sensitive (Table 9 ).
Figure 3: PL identity access despite similar pooled ranks: oracle (a) and predicted (b) lengths, 40,000 targets per query. Checkpoints and reconstruction: Table 1 .
Model
Eff. rank
R1 ↑
CER ↓
EM ↑
Pool
Token
Pool
Rel.
PL0
15.76
15.76
6.15
6.31
87.50
0.00
PL1
4.79
9.72
86.76
97.34
21.38
12.81
PL2
15.82
33.04
94.45
98.80
23.09
12.64
PL3
36.06
65.11
97.57
99.19
22.51
12.53
PL4
83.59
126.10
91.31
99.30
21.63
12.80
Table 1: All evaluated PL setting checkpoints at natural severity 3. Rank and reconstruction use oracle lengths; retrieval uses predicted lengths. R1, CER, and exact match (EM) are percentages.
Model
R1 (%) ↑
MRR ↑
WS3 (%) ↑
PL0
23.34
0.281
1.96
PL2
88.82
0.908
83.82
PL3
97.37
0.979
81.20
ByT5 encoder
87.48
0.894
22.05
SimCSE
98.71
0.991
98.71
NegMPNet
99.12
0.994
99.12
Table 2: LexNorm identity retrieval (1,941 queries, 1,954 targets). R1: original Recall@1; WS3: strongest whitespace perturbation. CANOPE uses predicted lengths and relative positions; external readouts are in Appendix B.6 .
Model
SAN ↑
PAWS ↑
STS-B ↑
SICK-R ↑
PL0
26.24
0.488
0.003
0.143
PL2
2.28
0.585
0.154
0.333
PL3
9.11
0.587
0.177
0.375
ByT5 encoder
8.12
0.584
0.248
0.283
SimCSE
63.80
0.647
0.842
0.808
NegMPNet
69.57
0.623
0.842
0.767
Table 6
System
Clean ↓
Augmented ↓
Δ
MatchaTTS
2.21
10.93
+8.71
Frozen PL2 + Matcha
13.60
21.54
+7.94
Frozen PL0 + Matcha
99.12
99.22
+0.09
Table 5: Downstream speech WER (% ↓ ); 3,930 utterances and 66,489 reference words per cell. Δ : augmented minus clean WER (percentage points), reported descriptively.
Appendix figures & tables39 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: The same increase in effective rank can help or hurt a fixed readout. (a) Increasing ρ transfers variance from the first direction to the second while keeping a+b=2 . (b) Ridge risk rises when the target uses the first direction and falls when it uses the second. The map remains invertible for every 0<ρ<1 , so both targets remain fully represented. These are exact population curves for d=2 , r=1 , μ=q0=1 , and zero observation noise, illustrating Theorem 5 .
Run
Predictive objective
Vectors regularized
CL-GP
Canonical-token grounding
Mean of canvas rows
CL-GT
Canonical-token grounding
One sampled row per item
CL-LP
Cross-view latent agreement
Mean of canvas rows
CL-LT
Cross-view latent agreement
One sampled row per item
Appendix
Table 6: CL setting design. All four use oracle target lengths and VISReg weight 0.01. Only G sends canonical-token loss gradients into the representation.
Run
Source ID
Defining setting
PL0/JEPA-based
a0_jepa_s42
Detached-anchor latent agreement
PL1/Baseline
a10_dec2_s42
Token supervision, ζ=0
PL2/VISReg-Weak
a2_visreg_s42
Pooled VISReg, ζ=0.1
PL3/VISReg-Stronger
a3_visreg_kuat_s42
Pooled VISReg, ζ=0.5
PL4/No-Invariance
a4_no_inv_s42
α=0
PL5/No-Negation
a5_no_neg_s42
β=0
Appendix
Table 7: Run names and their source identities. Short IDs in the text refer to this convention; raw result files retain the original checkpoint keys.
Model
Pooled rank
Token rank
Oracle R1 ↑
Predicted R1 ↑
CER ↓
PL0/JEPA-based
15.76
15.76
13.46
6.31
87.50
PL2/VISReg-Weak
15.82
33.04
98.80
98.80
23.09
PL3/VISReg-Stronger
36.06
65.11
99.16
99.19
22.51
Appendix
Table 8: PL setting at natural severity 3. Rank and CER use oracle routing; Recall@1 (R1) uses relative positions. Percentages except rank.
Model
Recov. oracle ↑
Recov. input ↑
Heldout oracle ↑
Heldout input ↑
CL-GP
99.985
97.475
84.278
84.665
CL-GT
99.980
16.532
89.472
31.220
CL-LP
0.520
0.090
8.183
6.759
CL-LT
0.747
0.002
0.715
0.014
Appendix
Table 9: CL setting relative-position Recall@1 (%) at severity 3. Input routing uses source length, not a learned length predictor.
Model
Channel
Or. CER ↓
Or. EM ↑
Or. CE ↓
In. CER ↓
In. EM ↑
CL-GP
recoverable
0.13
98.97
0.0094
91.84
0.09
CL-GP
heldout
18.13
6.86
16.3316
17.62
6.83
CL-GT
recoverable
0.15
98.78
0.0109
91.70
0.09
CL-GT
heldout
19.08
6.86
12.8150
17.06
6.83
Appendix
Table 10: Native-head reconstruction at E severity 3. CER and exact match (EM) are percentages; CE is corpus nats per target byte under oracle routing. CL-LP/CL-LT heads were not trained and are excluded.
Model
States
Eff. rank
∥μ∥↓
vˉ
DΣ↓
DCF↓
CL-GP
Pool
55.08
6.94
0.73763
2.456
0.03175
CL-GP
Token
41.59
6.94
26.37623
109.941
0.15207
CL-GT
Pool
114.14
1.04
0.02654
0.977
0.40322
CL-GT
Token
194.47
1.05
0.82981
1.375
0.00280
CL-LP
Pool
5.36
32.53
0.04498
1.103
0.69940
CL-LP
Token
15.57
32.53
0.06938
1.101
0.65024
Appendix
Table 11: Raw CL geometry at recoverable severity 3, oracle routing. vˉ is mean coordinate variance, DΣ=∥Σ−I∥F/d , and DCF is Eq. ( 44 ). Arrows indicate closer agreement with the zero-mean, identity-covariance Gaussian reference, not better content access; vˉ has target 1. Token statistics use one valid token per sentence.
Figure 5: Pooled effective rank and content access. The left panel uses PL natural severity 3 with oracle lengths for the principal configurations. The right panel shows that mean subtraction restores CL-LT clean identity access much more than cross-view retrieval. Both panels summarize full-cohort measurements.
Figure 6: Matched CL identities in a 2×2 grid. Each panel uses pooled states, oracle lengths, and its own clean PCA basis. Colors identify 12 matched texts; markers distinguish clean and recoverable severities 1–3. Lines connect views rather than time steps. CL-LP exhibits visible displacement for several identities, while some concentrated representations show short trajectories despite poor full-gallery retrieval. The figure illustrates selected within-model behavior, not an estimate of retrieval accuracy.
Figure 7: Relative-position retrieval across corruption strengths. The panels use PL predicted lengths, CL input lengths, and LexNorm changed-input queries with extra whitespace. Level 0 denotes clean PL/CL inputs and the original noisy LexNorm query. Channels and cohorts differ, so severity scales are not directly comparable.
Figure 8: Centering restores clean access, not robustness. CL-LT, oracle lengths, relative-position scoring, 40,000 targets. Blue subtracts the clean mean of 1,024 disjoint calibration examples; red uses raw states. Labels give Recall@1 (%).
Figure 9: Position improves access within a fixed checkpoint and predicted length. Gray pools states; blue compares 16 relative positions. All 1,941 changed-input queries search the same 1,954 targets. At maximum whitespace perturbation, PL2/PL3 gain 36.89/21.54 points.
System
Input
S ↓
D ↓
I ↓
WER ↓ [95% interval]
Zero-WER clips ↑
MatchaTTS
Clean
868
259
345
2.21 [2.04, 2.40]
3,083
MatchaTTS
Augmented
4,910
701
1,654
10.93 [10.47, 11.40]
1,647
Frozen PL2
Clean
6,657
1,398
988
13.60 [13.21, 14.01]
816
Frozen PL2
Augmented
10,880
2,085
1,359
21.54 [21.01, 22.09]
458
Frozen PL0
Clean
39,958
24,689
1,259
99.12 [98.33, 100.10]
0
Frozen PL0
Augmented
39,465
25,231
1,272
99.22 [98.45, 100.15]
0
Appendix
Table 12: Audited downstream error counts and utterance-paired percentile intervals. S/D/I denote substitutions, deletions, and insertions. The word denominator is 66,489 in every row.
Figure 10: Following the same identities through natural corruption. Each panel shows 12 texts, one color per text. Follow a line from the clean circle through the diamond, square, and cross at corruption levels 1, 2, and 3. The displayed states are averaged text vectors with true target lengths supplied. Axes are fitted on clean states separately for each model. Some PL0 views stay close even though full-gallery retrieval is poor in Figure 3 ; visual proximity must therefore be read together with identity access. Percentages state how much clean variance the three displayed axes retain.
Figure 11: What the geometry diagnostic sees in the four CL configurations. Rows identify the models; columns show clean text, reversible corruption at level 3, and an isotropic Gaussian reference. Each point is one averaged text vector; true target lengths are supplied. The same clean PCA axes and scale apply across a row. CL-GP/GT spread most variance beyond the three displayed axes, whereas CL-LP/LT concentrate it in this subspace. Compare these clouds with retrieval in Figure 2 , rather than reading a compact cloud as evidence of content preservation. Axes are fitted to 4,096 clean examples; each cloud displays 1,200 points. Surfaces span two standard deviations, not confidence regions.
Model
Representation
Effective rank
d90
d95
PC1–3 (%)
PL0
pooled
15.62
15
18
44.43
PL0
token
15.63
15
18
44.43
PL2
pooled
15.54
12
14
34.94
PL2
token
34.22
31
48
25.55
PL3
pooled
35.49
29
37
21.02
PL3
token
65.41
58
85
15.44
Appendix
Table 13: Clean 4,096-item visualization subset, oracle length. The dimensions d90 and d95 are the smallest numbers of principal components explaining 90% and 95% of variance. These subset ranks need not equal the full-cohort ranks.
Diagnostic
PL
CL
Total
Full-cohort summary charts
5
10
15
Full-cohort covariance spectra
2
3
5
Clean-subset spectra
6
8
14
3D distributions
6
16
22
Paired 3D trajectories
6
16
22
Projected Gaussian Q–Q plots
6
16
22
Appendix
Table 14: Coverage of the complete accompanying visual atlas. Each registered source PNG appears on its own landscape page. The panel index supplies page numbers, original paths, notes, and SHA-256 hashes.
Figure 12: PL natural distributions of pooled states with oracle length. Each row shows clean, severity 3, and the isotropic Gaussian reference in the same clean PCA coordinates. All supplied models are included. Clouds contain 1,200 points; surfaces span two standard deviations and are not confidence regions.
Figure 13: PL natural distributions of token states with oracle length. Each row shows clean, severity 3, and the isotropic Gaussian reference in the same clean PCA coordinates. All supplied models are included. Clouds contain 1,200 points; surfaces span two standard deviations and are not confidence regions.
Figure 14: Paired PL natural trajectories for token states with oracle length. Each panel follows 12 canonical identities through clean and severity levels 1–3 using one clean PCA basis per model. Colors identify targets and markers identify corruption levels. Differences in axes and captured variance require within-model interpretation.
Figure 15: Full-cohort PL natural diagnostics across severity, using oracle length. Retrieval, pooled and token effective rank, mean magnitude, variance scale, and cross-view error are shown separately. The bottom row uses a symmetric logarithmic scale with a linear region near zero. These curves use the original 40,000-item numeric summaries, independently of the smaller visualization subset.
Figure 16: CL recoverable distributions of token states with oracle length. Each row shows clean, severity 3, and the isotropic Gaussian reference in the same clean PCA coordinates. All supplied models are included. Clouds contain 1,200 points; surfaces span two standard deviations and are not confidence regions.
Figure 17: Paired CL recoverable trajectories for token states with oracle length. Each panel follows 12 canonical identities through clean and severity levels 1–3 using one clean PCA basis per model. Colors identify targets and markers identify corruption levels. Differences in axes and captured variance require within-model interpretation.
Figure 18: Full-cohort CL recoverable diagnostics across severity, using oracle length. Retrieval, pooled and token effective rank, mean magnitude, variance scale, and cross-view error are shown separately. The bottom row uses a symmetric logarithmic scale with a linear region near zero. These curves use the original 40,000-item numeric summaries, independently of the smaller visualization subset.
Figure 19: CL heldout distributions of pooled states with oracle length. Each row shows clean, severity 3, and the isotropic Gaussian reference in the same clean PCA coordinates. All supplied models are included. Clouds contain 1,200 points; surfaces span two standard deviations and are not confidence regions.
Figure 20: Paired CL heldout trajectories for pooled states with oracle length. Each panel follows 12 canonical identities through clean and severity levels 1–3 using one clean PCA basis per model. Colors identify targets and markers identify corruption levels. Differences in axes and captured variance require within-model interpretation.
Figure 21: CL heldout distributions of token states with oracle length. Each row shows clean, severity 3, and the isotropic Gaussian reference in the same clean PCA coordinates. All supplied models are included. Clouds contain 1,200 points; surfaces span two standard deviations and are not confidence regions.
Figure 22: Paired CL heldout trajectories for token states with oracle length. Each panel follows 12 canonical identities through clean and severity levels 1–3 using one clean PCA basis per model. Colors identify targets and markers identify corruption levels. Differences in axes and captured variance require within-model interpretation.
Figure 23: Full-cohort CL heldout diagnostics across severity, using oracle length. Retrieval, pooled and token effective rank, mean magnitude, variance scale, and cross-view error are shown separately. The bottom row uses a symmetric logarithmic scale with a linear region near zero. These curves use the original 40,000-item numeric summaries, independently of the smaller visualization subset.
Figure 24: LJ033-0072: all three systems under both input views. WER uses the clean reference; gray after a panel ends indicates no generated audio. Axes are shared within this example; the color reference is shared across all ten examples.
Figure 25: LJ006-0145: all three systems under both input views. WER uses the clean reference; gray after a panel ends indicates no generated audio. Axes are shared within this example; the color reference is shared across all ten examples.
Figure 26: LJ002-0239: all three systems under both input views. WER uses the clean reference; gray after a panel ends indicates no generated audio. Axes are shared within this example; the color reference is shared across all ten examples.
Figure 27: LJ038-0296: all three systems under both input views. WER uses the clean reference; gray after a panel ends indicates no generated audio. Axes are shared within this example; the color reference is shared across all ten examples.
Figure 28: LJ014-0142: all three systems under both input views. WER uses the clean reference; gray after a panel ends indicates no generated audio. Axes are shared within this example; the color reference is shared across all ten examples.
Figure 29: LJ012-0270: all three systems under both input views. WER uses the clean reference; gray after a panel ends indicates no generated audio. Axes are shared within this example; the color reference is shared across all ten examples.
Figure 30: LJ011-0213: all three systems under both input views. WER uses the clean reference; gray after a panel ends indicates no generated audio. Axes are shared within this example; the color reference is shared across all ten examples.
Figure 31: LJ007-0211: all three systems under both input views. WER uses the clean reference; gray after a panel ends indicates no generated audio. Axes are shared within this example; the color reference is shared across all ten examples.
Figure 32: LJ038-0221: all three systems under both input views. WER uses the clean reference; gray after a panel ends indicates no generated audio. Axes are shared within this example; the color reference is shared across all ten examples.
Figure 33: LJ006-0020: all three systems under both input views. WER uses the clean reference; gray after a panel ends indicates no generated audio. Axes are shared within this example; the color reference is shared across all ten examples.
Center for AI Research (CAIR), VinUniversity, Hanoi, Vietnam · Faculty of Computer Science and Engineering Ho Chi Minh City University of Technology (HCMUT), VNU-HCM Ho Chi Minh City, Vietnam · Mohamed bin Zayed University of Artificial Intelligence Abu Dhabi, United Arab Emirates