cs.LGOct 6, 2026

CNet: A Complex-Valued Deep Learning Framework with Wirtinger Autodifferentiation and FFT--Hadamard Convolution

Authors: Marcel Crasmaru

Abstract

CNet is a C++/CUDA framework for building deep complex-valued neural networks (CVNNs) and optimizing complex functions by gradient descent with Wirtinger derivatives. Complex models are underexplored yet natural where data is intrinsically complex -- RF/IQ communications, MRI k-space, radar/SAR, audio spectra -- and phase carries information real networks discard. CNet is physics-native: a network is a cascade of complex (often unitary) operations on an amplitude vector, and classification is a Born-rule measurement p_k = |z_k|^2/||z||^2 rather than a softmax over real logits. Its library of complex layers makes conv(x,k) = IFFT(FFT(x).FFT(k)) learnable via signal-processing primitives -- spectral-padding local kernels (Pad), the inverse DFT, a magnitude nonlinearity |z|^2 (CModulus2), holomorphic powers z^M (CPower), and mean pooling (MeanPool) -- with GPU kernels for every layer, a GPU true-Adam optimizer, and a reduced-memory inference mode. Three studies. (1) A fully complex FNet-style causal language model, built on a new O(N log N) causal Fourier mixer, matches a param-matched real-valued causal FNet on tiny-shakespeare in under half the steps. We introduce Born-rule attention: a content-based causal mixer whose weights are quantum-measurement probabilities |<Q_k,K_j>|^2 of one token against another -- the first attention mechanism built on the Born rule. With rotary position embeddings and per-token complex normalization, a fully complex Born-rule model reaches 1.52 nats/char on tiny-shakespeare, matching softmax attention and beating the Fourier mixer by 0.23. (2,3) Two bottleneck analyses -- radio-modulation classification (RML2016.10a) and the Fourier phase problem of coherent-diffraction imaging -- isolate where CVNNs need new operators. The complex machinery learns the correct structure; open problems are concrete and operator-level. Code: https://github.com/crasmarum/CNet

Figures & tables

Explore similar work

May 26, 2026cs.LG

When do complex-valued neural networks help? A study of representation, geometry, and optimization

Complex-valued Neural Networks (CVNNs) are often motivated by domains where information is naturally encoded in magnitude and phase. Yet complex-valued inputs alone do not determine when complex arithmetic improves learning: the label signal may lie in amplitude, phase, their coupling, or a symmetry that real-valued models can also represent under suitable coordinates. We study this through a representation-first evaluation of CVNNs against Cartesian real, polar, phase-only, magnitude-only, parameter-matched real, and FLOP-matched real baselines. Across synthetic RF tasks, complex representations are useful but not universally superior. PSK-only tasks favor phase-aware and complex-valued models, QAM-only tasks favor magnitude-based models, mixed PSK+QAM gives only a small complex-valued advantage, and unseen carrier-phase rotations break coordinate-dependent models without augmentation. Similar patterns appear beyond RF: in quantum-wavefunction prediction, momentum is invisible to ∣ψ∣|ψ| but recoverable from phase, while EEG analytic-signal experiments show that phase locking, amplitude bursts, and phase-amplitude coupling each favor different coordinate views. We also identify a benchmarking artifact on RadioML 2018.01A. Under matched-shared-trial selection, a CReLU complex model exceeds the best real baseline by 22.94 PP; under independent per-family tuning on the same data and 16-trial search space, the gap collapses to 2.46 PP. Gradient analysis traces the inflated gap to high-learning-rate first-step instability in real baselines, while complex parameter coupling distributes the loss signal more robustly. A learning-rate ×\times activation factorial confirms the failure is primarily hyperparameter-driven. Overall, CVNNs are best viewed as structured inductive biases whose gains depend on representation, symmetry, and optimization, not as universally superior architectures.
Apr 21, 2026cs.AR

Algorithm and Hardware Co-Design for Efficient Complex-Valued Uncertainty Estimation

Complex-Valued Neural Networks (CVNNs) have significant advantages in handling tasks that involve complex numbers. However, existing CVNNs are unable to quantify predictive uncertainty. We propose, for the first time, dropout-based Bayesian Complex-Valued Neural Networks (BayesCVNNs) to enable uncertainty quantification for complex-valued applications, exhibiting broad applicability and efficiency for hardware implementation due to modularity. Furthermore, as the dual-part nature of complex values significantly broadens the design space and enables novel configurations based on layer-mixing and part-mixing, we introduce an automated search approach to effectively identify optimal configurations for both real and imaginary components. To facilitate deployment, we present a framework that generates customized FPGA-based accelerators for BayesCVNNs, leveraging a set of optimized building blocks. Experiments demonstrate the best configuration can be effectively found via the automated search, attaining higher performance with lower hardware costs compared with manually crafted models. The optimized accelerators achieve approximately 4.5x and 13x speedups on different models with less than 10% power consumption compared to GPU implementations, and outperform existing work in both algorithm and hardware aspects. Our code is publicly available at: https://github.com/zehuanzhang/BayesCVNN.git.
Feb 6, 2026cs.LG

Perturbing the Phase: Analyzing Adversarial Robustness of Complex-Valued Neural Networks

Complex-valued neural networks (CVNNs) are rising in popularity for all kinds of applications. To safely use CVNNs in practice, analyzing their robustness against outliers is crucial. One well known technique to understand the behavior of deep neural networks is to investigate their behavior under adversarial attacks, which can be seen as worst case minimal perturbations. We design Phase Attacks, a kind of attack specifically targeting the phase information of complex-valued inputs. Additionally, we derive complex-valued versions of commonly used adversarial attacks. We show that in some scenarios CVNNs are more robust than RVNNs and that both are very susceptible to phase changes with the Phase Attacks decreasing the model performance more, than equally strong regular attacks, which can attack both phase and magnitude.