Curvature Under Attack in hZACH-ViT: Gauge Symmetry, Boundary Saturation, and Adversarial Failure
Authors: Athanasios Angelakis, Marta Gomez-Barrero
Organizations: BioML, Research Institute CODE, University of the Bundeswehr Munich, Munich, Germany · Amsterdam UMC, University of Amsterdam, Amsterdam, Netherlands
Curvature is often treated as an intrinsic property of a representation, although its empirical effect also depends on coordinate scale, learned logit temperature, and numerical safeguards. We study this interaction in hZACH-ViT, a compact Vision Transformer with Euclidean, Poincare, and spherical prototype heads. The backbone architecture, seed-specific initialization, 50-per-class training subset, and optimization protocol are matched across three MedMNIST datasets and five seeds. At the fixed comparison curvature c=1, Poincare has the lowest class-macro PGD attack-success rate in all 12 dataset-budget cells and under a stronger CE+DLR multi-restart attack on all three datasets, but it also has the lowest clean MacroF1. An end-to-end curvature intervention changes the interpretation. Reducing Poincare curvature to c=0.1 improves clean MacroF1 in every one of the 15 paired seed-dataset comparisons and removes hard boundary clipping, yet on OrganAMNIST it increases strong attack success from 89.7% to 99.3% (paired difference +9.57 points; 95% hierarchical bootstrap CI [+5.52,+14.03]). At c=1, 40-47% of clean Poincare features are hard-clipped, the radial Jacobian of the inherited map is nearly zero, and dimensionless attack trajectories are unusually long and inefficient. The spherical head provides a control: its curvature change is an exact scale gauge to floating-point precision and produces much smaller attack differences. These results do not establish intrinsic hyperbolic robustness. They identify an implementation-sensitive regime in which curvature, scale, and proximity to the Poincare boundary jointly organize clean recognition and adversarial representation motion.
Figures & tables
Figure 1: PGD-10 class-macro ASR for the fixed c=1 comparison. Curves are five-seed means and shaded regions are hierarchical 95% bootstrap intervals. Every head is attacked on the identical four-head shared-clean-correct subset within each seed. Lower ASR means slower empirical failure under this attack, not certified robustness.
Dataset
Head
Clean F1
PGD-10
PGD-20/R3 CE+DLR
Blood
Linear
.805
97.70
96.88
Euclidean prototype
.783
95.39
96.41
Poincaré, c=1
.714
93.79
93.59
Spherical, c=1
.770
94.73
94.84
Derma
Linear
.354
97.31
98.33
Euclidean prototype
.340
95.36
97.71
Table 1: Clean MacroF1 and ASR (%) at 4/255 , averaged over five seeds. PGD-10 and strong attacks use independently capped versions of the same four-head shared-clean-correct pool. Bold marks the clean-best or attack-slowest head within dataset.
Figure 2: End-to-end curvature intervention. Small points are seeds, large points/diamonds are five-seed means, and lines in (a-c) connect matched seeds. (a) Poincaré clean MacroF1. (b) Strong class-macro ASR. (c) Fraction of clean Poincaré features reaching the hard projection threshold. (d) Change in dimensionless PGD-10 path efficiency at 4/255 ; positive values indicate a more direct trajectory at c=0.1 . The spherical family is the curvature-gauge control.
Dataset
Clean MacroF1
Strong ASR (%)
Median ρc
Hard clip (%)
Blood
.714 → .791
93.91 → 94.84
1.000 → .981
40.1 → 0
Derma
.317 → .350
97.71 → 98.33
1.000 → .972
47.2 → 0
Organ
.405 → .504
89.70 → 99.27
1.000 → .943
44.2 → 0
Table 2: Poincaré end-to-end curvature sensitivity. Radius and clipping entries are averages of seed-level clean diagnostics. Arrows run from c=1 to c=0.1 .
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Family
Direction
Max ∣Δℓ∣
Mean ∣Δℓ∣
Pred. changes
Logits
Sphere
1→0.1
4.77×10−7
7.27×10−8
0
66,621
Sphere
0.1→1
1.43×10−6
2.06×10−7
0
66,621
Poincaré
1→0.1
7.13×10−1
7.46×10−3
1
62,572
Poincaré
0.1→1
3.99×10−1
5.52×10−4
0
62,572
Appendix
Table 3: Curvature-scale gauge audit. Mean errors are weighted by the number of logits.
Figure 3: OrganAMNIST Poincaré convergence under the first restart of the strong attack, averaged over five seeds. CE and DLR strengthen through 20 steps while preserving the c=1 versus c=0.1 separation. Final ASR uses the worst true-class margin over all six CE/DLR restart candidates.
Dataset
Configuration
Endpoint
Path
Efficiency
Clip steps
Blood
Poincaré c=1
24.33
205.70
.120
.404
Poincaré c=0.1
8.67
57.82
.155
.000
Derma
Poincaré c=1
24.14
199.55
.123
.449
Poincaré c=0.1
7.71
41.31
.207
.000
Organ
Poincaré c=1
23.97
215.27
.118
.579
Poincaré c=0.1
7.38
58.24
.211
.000
Appendix
Table 4: Five-seed mean PGD-10 trajectory quantities at 4/255 . Path and endpoint use cdc . “Clip steps” is the fraction of trajectory states at the Poincaré hard projection threshold.
Compact Vision Transformers are attractive for medical imaging in low-data and resource-constrained settings, but most existing variants assume that Euclidean latent geometry is sufficient for organizing image representations. We introduce hZACH-ViT, a family of curved-geometry extensions of ZACH-ViT, a compact zero-token Vision Transformer that removes positional embeddings and the class token and relies on global average pooling over patch representations. To isolate the role of geometry, we preserve the verified ZACH-ViT backbone and modify only the final representation space and prototype-based classifier head, enabling a controlled comparison between Euclidean, hyperbolic, and spherical latent geometries. We evaluate Poincaré, Klein, and spherical hZACH-ViT heads on seven MedMNIST datasets under an identical few-shot protocol with 50 samples per class and five random seeds. The completed benchmark contains 770 training runs spanning seven datasets, three non-Euclidean geometries, seven curvature magnitudes, and a Euclidean baseline. Across all seven datasets, the best non-Euclidean hZACH-ViT configuration improves over Euclidean ZACH-ViT, with an average gain of +0.021 in the dataset-specific primary metric and the largest improvement on OCTMNIST (+0.055 MacroF1). Fixed low-curvature configurations retain positive gains on the majority of datasets, and low curvature values (c = 0.1 or 0.2) account for six of the seven dataset-level winners. Rather than identifying a universally optimal manifold, our results establish geometry and curvature as dataset-dependent model-selection variables, with fixed low-curvature analyses confirming that gains persist beyond exhaustive per-dataset tuning.
Athanasios Angelakis
BioML Lab, Research Institute CODE, UniBw, Munich, Germany · Department of Epidemiology and Data Science, Amsterdam UMC, Amsterdam, Netherlands
Whether a hyperbolic representation model uses its geometry cannot be inferred from curvature alone: what matters is the dimensionless operating point cρ and whether the radial and cone mechanisms are operational there. We develop necessary-condition diagnostics and audit three published hyperbolic vision-language families -- MERU, HyCoCLIP, and PHyCLIP -- across released checkpoints and matched interventions. All converged checkpoints remain near-Euclidean (H(u)≈1; none reaches cρ>1), and releasing the curvature floor changes c and norms without leaving this regime or substantially degrading downstream performance. Entailment cones are inactive or saturated, and graded traversal fails under controlled readouts, including the models' native distance metrics. External parent-child ordering shows no shuffle-controlled pair-specific radial signal at quantified sensitivity; the only surviving pair-specific signal, a statistically detectable but small residual on the GRIT box-to-full-caption relation, remains non-operative under the evaluated readouts. Taxonomy correlations show no detectable norm contribution beyond cosine, and coarse-retrieval gains co-vary with box/compositional supervision without establishing an active radial mechanism. Gradient diagnostics expose a low-curvature, wide-cone shortcut in the entailment objective. A closed-form aperture identity places the saturation edge at cρ≤2K: with the floor released, all trained relation-level parent means lie at or below this edge, leaving the parent cones fully or nearly saturated. Entailment-off runs pass the edge and continue contracting. The shortcut is the dominant accelerator of collapse, not its sole cause. These audited formulations do not show an operative radial/cone mechanism under our diagnostics. We distill the audit into a five-number geometry report for hierarchy claims.
Training billion-parameter Transformers is often brittle, with transient loss spikes and divergence that waste compute. Even though the recently developed Edge of Stability (EoS) theory provides a powerful tool to understand and control the stability of optimization methods via the (preconditioned) curvature, these curvature-controlling methods are not popular in large-scale Transformer training due to the complexity of curvature estimation. To this end, we first introduce a fast online estimator of the largest (preconditioned) Hessian eigenvalue (i.e., curvature) based on a warm-started variant for power iteration with Hessian-vector products. We show theoretically, and verify empirically, that the proposed method makes per-iteration curvature tracking feasible at billion parameter scale while being more accurate. Using this tool, we find that training instabilities coincide with surges in preconditioned curvature and that curvature grows with depth. Motivated by these observations, we propose architecture warm-up: progressively growing network depth to carefully control the preconditioned Hessian and stabilize training. Experiments on large Transformers validate that our approach enables efficient curvature tracking and reduces instabilities compared to existing state-of-the-art stabilization techniques without slowing down convergence.
Sameera Ramasinghe, Ajanthan Thalaiyasingam, Hadi Mohaghegh Dolatabadi +6