Organizations: Department of Computer Science and AI Bar-Ilan University Ramat-Gan, Israel · Electrical and Computer Engineering Technion Haifa, Israel
Canonical Correlation Analysis (CCA) is a fundamental method for multiview shared space learning. However, its strict reliance on paired data poses a significant limitation, as such data is often difficult to obtain or entirely unavailable. In this paper, we present Unpaired CCA (UCCA), a novel method that learns linear projections to maximize the correlation of the true underlying pairing without access to any paired samples during training. We first establish theoretical results connecting the Quadratic Assignment Problem (QAP) to CCA. Leveraging these theoretical insights, we derive a practical method to maximize correlation exclusively from unpaired data. To the best of our knowledge, UCCA is the first approach to learn maximally correlated projections in a strictly unpaired setting. We validate UCCA on real-world multi-modal datasets, demonstrating that it significantly outperforms recent unpaired alignment baselines in recovering the underlying true correlation. This work fills a critical gap between traditional statistical multiview learning and the growing field of unpaired data learning.
Figures & tables
Figure 1: The theoretical and empirical alignment of Unpaired CCA and QAP. (a) QAP vs. CCA: On a 10-sample synthetic dataset, evaluating all 10! possible pairings reveals a direct positive relationship between the QAP value and total correlation (App. E.1 ). (b) UCCA Performance: On the Handwritten benchmark (Sec. 6 ), UCCA captures high correlation despite training fully unpaired, significantly outperforming unpaired alignment baselines and approaching the paired CCA upper bound.
Figure 2: MCP and MCPO(d) demonstration. (a) MCP successfully reconstructs the true pairing on aligned data. (b) MCP depends on input rotation, which leads to poor performance on unaligned data. (c) Conversely, MCPO(d) optimizes over all orthogonal alignments, choosing the pairing with the highest correlation after the projection. In this example, it chooses the pairing from (a), enabling it to retrieve the true pairing, invariant to input (orthogonal) projections. A dynamic animation of this figure is available on our project page .
Figure 3: (a) Validation of Assum. 1 : Partial cumulative sums of singular values for the true permutation P∗ (#shuffles = 0) and various random shuffles on the Flickr dataset (8,000 samples; see Sec. 6 ). As the permutation diverges from P∗ , the partial sums strictly decrease, confirming the majorization property. (b) Validation of Thm. 2 . A comparison of MCPO(d) (Nuclear) and MCPSF (Frobenius) objectives using the setup from (a). Both objectives decay as the permutation moves away from P∗ , and both are locally maximized by the same ground-truth permutation.
Figure 4
Figure 5: Total Correlation results. Total correlation achieved across six distinct dataset configurations. UCCA consistently extracts highly correlated components, significantly outperforming all baselines, and often approaching the paired CCA upper bound.
Figure 6: MCPO(d)=MCPSF in practice. Kendall Tau Distances between MCPSF and approximated MCPO(d) , computed on the anchors of the Handwritten benchmark (PC-PA) across 100 initializations. The strict concentration at zero empirically validates Thm. 2 .
Figure 7: Runtime and Performance. Total correlation versus runtime for UCCA and baselines on the handwritten (PC-PA) benchmark. UCCA is improved in robustness and accuracy with more clustering repetitions, yet outperforms all baselines even with fewer repetitions and shorter runtimes.
Figure 8: Total correlation performance on the Handwritten benchmark (PC-PA) across (a) data imbalances and (b) varying amounts of unpaired data . UCCA consistently outperforms baselines, achieving its peak performance when the maximum amount of data is accessible. This demonstrates that abundant unpaired data alone provides a strong alignment signal, improving performance without requiring any paired correspondences.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Extended validation of MCPO(d)=MCPSF . Distance between MCPSF and the approximated MCPO(d) (denoted as P~O ), computed on the anchors of the each dataset across 100 initializations. This figure complements Fig. 6 in the main text by demonstrating that the strict concentration at zero holds consistently across all evaluated domains, further validating Thm. 2 .
Figure 10: Extended runtime and performance analysis. Total correlation versus runtime (in seconds, log scale) for UCCA and baseline methods across the various datasets. The green numerical annotations on the UCCA curve denote the number of clustering repetitions. UCCA rapidly surpasses baseline performance even at a lower repetition setting and improves as repetitions increase, while maintaining a practical total execution time.
Figure 11: Effect of the QAP solver on performance. Total Correlation achieved across multiple initializations on the PC-PA Handwritten benchmark, using the 2-Opt and FAQ approximate QAP solvers. Both solvers yield very similar overall performance and alignment quality. Because FAQ failed to converge in two instances (points near zero), 2-Opt was selected as the default solver for UCCA simply to ensure consistent execution across all runs.
Figure 12: Effect of the number of anchors. Total Correlation versus the number of anchors used in the matching step. Increasing the number of anchors allows UCCA to capture the manifold geometry more accurately, leading to a sharp performance improvement that approaches the paired CCA upper bound (represented by the dashed horizontal line). The anchor count 20 is chosen to balance this high accuracy with computational efficiency.
Figure 13: Effect of the number of PCA dimensions. Total Correlation versus the number of PCA components used in the preprocessing step.
Figure 14: Embeddings visualizations. Visualizations of the generated embeddings of Paired-CCA and UCCA.
Figure 15: Synthetic toy dataset.
Figure 16: Extended empirical validation of Assum. 1 . Partial cumulative sums of singular values for the true permutation ( P∗ , denoted by 0 shuffles) and various randomly shuffled permutations across all datasets: PC-KL, PC-PA, KL-PA, SNARE, Flickr, and COCO. Consistent with the main text results (Fig. 3 a), the partial sums strictly decrease as the permutation diverges from P∗ , verifying the weak majorization property across all tasks.
Figure 17: Extended empirical validation of Theorem 2. Comparison of the MCPO(d) (Nuclear Norm) and MCPSF (Squared Frobenius Norm) objectives as the permutation is randomized away from the true alignment P∗ across all datasets: PC-KL, PC-PA, KL-PA, SNARE, Flickr, and COCO. In alignment with the main text results (Fig. 3 b), both objectives decay synchronously as the number of random shuffles increases, empirically confirming that both formulations are locally maximized by the same ground-truth permutation.
Figure 18: Data preparation: While the test is kept paired for evaluation purposes, we split the train such that no correspondence is available during training. This exactly replicates the unpaired settings discussed in the paper.
We investigate the identifiability of nonlinear canonical correlation analysis (CCA) in a multi-view setup, in which each view is generated by applying an unknown nonlinear map to a linear mixture of shared latent variables plus view-private noise. Rather than pursuing exact unmixing, which is known to be ill-posed under general nonlinear mixing, we instead reframe multi-view CCA as a basis-invariant subspace identification problem. Under suitable latent priors and spectral separation conditions, we prove that the pairwise population CCA objective recovers correlated signal subspaces up to view-wise orthogonal ambiguity. For N≥3 views, their multi-view aggregation provably isolates the jointly correlated subspaces shared across all views while eliminating view-private variation. We further establish finite-sample statistical consistency guarantees by translating the concentration of empirical cross-covariances into explicit subspace error bounds via spectral perturbation theory. Experiments on synthetic and rendered image datasets support our theoretical findings and illustrate the necessity of the assumed conditions.
Zhiwei Han, Stefan Matthes, Hao Shen
Chair of Data Processing, Technical University of Munich, Germany · fortiss GmbH, Munich, Germany
Gradient-boosted trees dominate tabular machine learning, yet canonical correlation analysis has always relied on linear or neural encoders. We propose \textbf{TreeCCA}, the first method to train gradient-boosted tree ensembles end-to-end as CCA encoders, inheriting their plug-and-play reliability: no architecture design, familiar hyperparameters, and strong performance with defaults. The technical enabler is the Eckart-Young (EY) loss, which supplies closed-form per-sample gradients that slot directly into any standard GBT library (XGBoost, LightGBM) as a custom objective. TreeCCA is the first CCA method to combine nonlinear accuracy with native interpretability: every tree split selects one feature, so gain importances reveal which inputs drive cross-view correlation at no extra cost. We demonstrate these properties on synthetic benchmarks, where TreeCCA matches or exceeds Deep CCA (2.61 vs.\ 2.43 on Signed Power; 2.93 vs.\ 2.89 on Hermite), and on a sparse benchmark with zero linear cross-view covariance, where TreeCCA recovers the true support with Precision@S=1.00 at p=50 while PMD finds no signal. On the UCI HAR sensor-fusion benchmark, TreeCCA achieves comparable accuracy to Deep CCA at 5× lower cost, while XGBoost gain importances directly validate a physics-motivated hypothesis about the data --- an interpretation not readily available with neural encoders. Across five popular tabular multi-view datasets, TreeMCCA consistently matches or exceeds linear CCA in both nonlinear correlation extraction and downstream classification accuracy.
Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds and when each fails --- a gap that leaves practitioners, especially in scientific domains with heterogeneous instruments and multiple levels of measurement, unable to diagnose why standard methods underperform the best single modality. We study both objectives under a spiked signal-plus-noise model with structured cross-modal nuisance correlation, the ingredient that breaks the classical recovery guarantees, and derive separation ratios that expose complementary failure modes: alignment whitens each modality and fails when nuisance is strongly correlated across views; prediction encodes whatever is cross-predictable through a one-sided whitening, with recovery governed by source-modality quality. The resulting phase diagram partitions multimodal problems into four regimes --- Both, CA only, CP only, and Neither --- refined by a recovery count that separates partial recovery from complete failure. We present a data-driven procedure to locate real-world datasets in this diagram using a small labeled subsample, identifying the preferred objective and prediction direction before any cross-modal training, and identifying when no objective in the CA/CP family can improve on the stronger modality alone. Experiments on synthetic data, stereo-vision benchmarks, image--caption pairs, and two real scientific domains --- astronomy and single-cell multi-omics --- validate the predictions in the nonlinear regime, including both faces of the Neither regime. Code to reproduce the results is available at https://github.com/IlayMalinyak/mm_align_vs_pred.