Organizations: Computer Science Research Institute, Sandia National Laboratories, Albuquerque, NM · Department of Statistics & Data Science, University of California, Los Angeles, CA
When minimizing the squared-error loss, the popular CP decomposition can be interpreted as parameter inference in a Gaussian model with a low-rank mean tensor and constant variance across the tensor entries. We introduce heteroskedastic-CP (HCP), which models entrywise variability with a non-constant, low-rank precision tensor, and develop an alternating block-coordinate ascent method to recover both the low-rank mean and precision tensors from noisy observations. Our procedure is computationally competitive, with the same leading-order factor-update complexity as CP-ALS. We demonstrate HCP on synthetic experiments and an EEG application.
Figures & tables
Figure 1
Figure 1 : Frontal slice of S following heteroskedastic variance construction, for different values of λ .
Figure 2 : Comparison of relative errors in estimates of B at varying heteroskedasticity levels.
Figure 3 : Visualization of 9 th slice of the noisy observed tensor X and the CP, PGCP, and HCP estimates B , for varying heteroskedasticity levels.
Figure 4 : Panels correspond to a different true observation model, and boxplots show the B recovery in relative Frobenius error over replicates for HCP, CP, and GCP methods using Poisson, Gamma, and inverse Gaussian losses.
Figure 5 : Simulation results comparing recovery performance across tensor sizes. The left panel shows relative error in recovery of B , while the right panel shows relative error in recovery of S .
Figure 6 : Estimated log-variance from HCP for the EEG tensor, averaged over channels and runs. The recovered variance is largest at lower frequencies.
Figure 7 : Visualization of the factor matrices recovered from the EEG tensor by CP and HCP with rows corresponding to components. Within each method, columns show the channel/scalp loadings, time-frequency loadings, and measurement loadings. Dashed vertical lines in measurement columns indicate recording-run boundaries.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8 : Results for a toy 3 mode tensor ( 10×8×12 ) with R=3 and known S . The first row displays the first horizontal slice, and the second, the first lateral slice, for the observed data, the true mean, and the estimated mean.
Figure 9 : Results for a 3 mode tensor ( 10×8×12 ) with R=3 and known B=0 . The first row displays the first horizontal slice, and the second, the first lateral slice, for the squared residuals, the true variance, and the estimated variance.
Figure 10 : Results for a 3 mode tensor ( 10×8×12 ) with R=3 . For the first frontal slice we show the data, true mean tensor, and estimated mean tensor in the first row, and true vs learned variance tensors in the second row.
Figure 11 : Simulation results comparing recovery performance across tensor sizes and heteroskedasticity structures. The left panel shows relative error in recovery of B , while the right panel shows relative error in recovery of S .
Figure 12 : Mean recovery when varying the fitted mean rank RB , with the precision rank fixed at its true value RS=3 . The true mean rank is RB=3 . HCP consistently outperforms CP decomposition, with the worst performance for both methods occuring when the fitted mean rank is too small.
Figure 13 : Precision recovery when varying the fitted mean rank RB , with the precision rank fixed at its true value RS=3 . This plot shows how misspecification of the mean rank affects estimation of the precision tensor.
Figure 14 : Mean recovery when varying the fitted precision rank RS , with the mean rank fixed at its true value RB=3 . Importantly, even for the worst fitted precision ranks, HCP will outperform CP in mean recovery.
Figure 15 : Precision recovery when varying the fitted precision rank RS , with the mean rank fixed at its true value RB=3 . The true precision rank is RS=3 . Unlike for the mean, if the tensor is small, the fitted rank being too large can be more detrimental than being too small.
Figure 16 : Visualization of the factor matrices recovered from the EEG tensor by HCP and CP with rows corresponding to components. Within each method, columns show the channel/scalp loadings, time-frequency loadings, and measurement loadings. Dashed vertical lines in measurement columns indicate recording-run boundaries.
Low-rank tensor decomposition (TD) is usually effective on clean, fully observed data, but it often degrades under severe missingness or noise. Low-rankness is itself a useful but limited structural prior, and additional handcrafted priors (e.g., sparsity or smoothness) still fall short of capturing the rich statistics of real-world data. To compensate for this weak inductive bias under heavy corruption, one would like to inject a learned, data-driven prior; however, the state-of-the-art diffusion models are not readily compatible with current TD and tractable posterior inference. To address these challenges, we introduce DiffBCP, a hybrid-prior Bayesian CP decomposition framework that couples a cumulative shrinkage process prior over the CP factors for automatic rank selection with an off-the-shelf pre-trained diffusion model as an implicit data prior on the reconstructed tensor. To make posterior inference tractable despite the coupling among the likelihood, low-rank constraint, and diffusion prior, we develop a split Gibbs sampler: CP factors admit conjugate updates, while the diffusion block is sampled via low-rank-guided denoising. A noise-adaptive coupling schedule further reduces sensitivity to hand-tuned annealing. Experiments on image inpainting and denoising, including high-resolution out-of-distribution images, show consistent gains over Bayesian, nonlinear, and plug-and-play TD baselines.
Zerui Tao, Qibin Zhao
RIKEN Center for Advanced Intelligence Project (AIP), Tokyo, Japan.
This paper proposes Spectra-Guided Neural Tucker Factorization (SG-NTF) for High-Dimensional and Incomplete (HDI) tensor completion. Circumventing discrete representational limits, SG-NTF maps scalar timestamps into a continuous spectral space to abstract temporal periodicities. Concurrently, a Spatio-Temporal Co-Gating (STCG) mechanism explicitly filters latent interactions via multiplicative modulation on spatiotemporal contexts. Evaluations on real-world HDI tensors verify that SG-NTF maintains competitive completion accuracy with parameter efficiency.
Fusheng Wang, Yikai Hou
School of Automation, Chongqing University of Posts and Telecommunications, Chongqing, China · College of Computer and Information Science, School of Software, Southwest University, Chongqing, China
High-dimensional incomplete (HDI) tensors are widely used in traffic and climate applications, but sparse observations make accurate completion difficult. The intrinsic non-linear dynamics and non-stationary variations across distinct multi-modal fields severely hinder the efficacy of conventional linear reconstruction frameworks. Neural Tucker factorization provides an effective framework for modeling high-order interactions among tensor modes. By parameterizing underlying structural characteristics into continuous latent spaces, neural representations circumvent the rigid low-rank constraints of classical algebra. However, its performance can still be affected by implementation-level choices, especially parameter initialization and the bias configuration of the final output mapping. Suboptimal initializations frequently lead to variance explosion across the cubically expanded interaction spaces, driving the subsequent non-linear activation boundaries into severe gradient saturation zones, while the omission of a dedicated translation parameter forces interaction weights to implicitly absorb global statistical deviations. This paper proposes a simple yet effective neural Tucker factorization model with Kaiming initialization and bias correction (KaBiN) for HDI tensor completion. The proposed model utilizes Kaiming uniform initialization for the embedding and Tucker linear parameters, and adopts a simple bias correction in output mapping. By elegantly decoupling global mean shifts from local structural representations, the framework provides a highly stable and well-conditioned optimization landscape. Experiments on three real-world HDI tensor datasets show that KaBiN achieves better performance than the original NeuTucF, while introducing minimal computational overhead.
Yuchao Su, Yixin Ran
School of Computer Science and Engineering, Chongqing University of Science and Technology, Chongqing, China · College of Computer and Information Science, School of Software, Southwest University, Chongqing, China