Low-Latency Activation-Regularized Sparse Neural Operators with Distillation Assistance Towards Real-Time Neuromorphic Virtual Sensing
Authors: William Howes, Farid Ahmed, Syed Bahauddin Alam
Organizations: Grainger College of Engineering, Nuclear, Plasma & Radiological Engineering Department, University of Illinois Urbana-Champaign, Urbana, IL, USA · National Center for Supercomputing Applications, Urbana, IL, USA
Virtual sensing enables digital twins and safety-critical systems to reconstruct and forecast spatial-temporal physics in real time. However, conventional computational and data-driven methods often face challenges in generalization, latency, and energy efficiency for edge deployment. Neural operators offer a promising alternative but remain reliant on power-intensive hardware. Spiking neurons and neuromorphic computing can improve efficiency, yet surrogate-gradient training and multi-step spiking introduce convergence and latency challenges. We propose the Sparse-Activation-ReLU (SAR) layer, a single-step alternative that promotes activation sparsity without surrogate-gradient training while remaining compatible with event-based computing. Within a trunk-based NOMAD architecture, SAR achieves over a fivefold improvement in the combined Latency-Error-Energy (LEE) metric compared with Variable Spiking Neuron (VSN) and Leaky Integrate-and-Fire (LIF) implementations. We further analyze spiking entropy and feature usage and introduce synthetic knowledge distillation, reducing the LEE score by more than twofold. Finally, we improve VSN through a ReLU-based spiking loss and graph-neighbor thresholding. On the Heat Exchanger dataset, these approaches reduce L2 error by more than twofold and nearly sevenfold, respectively, while reducing spiking and spatial aggregation. Overall, the work presented is a step towards energy-efficient virtual sensing by providing an alternative framework that can be positioned towards neuromorphic or other edge device integration that can be a gold standard to compare latency, energy, and error performance for future efficient designs that are sparsity or brain-inspired spiking based.
Figures & tables
Figure 1: Sparse-Activation-ReLU Model (a) Visual demonstration of the Variable Spiking Neuron [ 17 ] operating after a linear fully-connected layer. Synaptic weights and unique bias terms accumulate the input for each neuron at spike time step t which is aggregated into the neuron’s leaky memory, mirroring LIF dynamics. VSN replaces the binary spike output with an activation function operating on the neuron input. (b) Sparse-Activation-ReLU for a fully-connected linear layer. We apply a ReLU activation function to the linear layer output which only allows positive signals subsequently operated on by a chosen activation function. Such forward pass communication allows for surrogate-free ANN-to-neuromorphic conversion that allows for variable signals (unlike traditional binary spike rate to activation matching), included in the original VSN framework. (c) Sparse-Activation-ReLU for a generic input (spectral/spatial convolution or norm). We apply a ReLU activation function to the input subtracted by an optional threshold utilized for better spiking control. Similar to the linear layers, the ReLU only allows positive signals subsequently operated on by a chosen activation function.
Figure 2: Sparse-Activation-ReLU Nonlinear Manifold Decoder for Operator Learning Architecture of SAR-NOMAD specifically for the 2D Heat Exchanger. In green, exists every SAR layer which naturally replaces the ReLU layers. Since every computation before SAR is a simple linear layer, we do not include a threshold parameter which is redundant in the presence of the bias term. As shown, the full-field output provides the L2 error for the total objective function while the individual sparse activation outputs of the ReLU (from all SAR layers in the branch, trunk and combined networks not fully depicted in the figure) with the Hoyer/ L1 loss term provide the control on energy efficiency.
Figure 3: Sparse-Activation-ReLU Graph Neural Operator Architecture of SAR-GNO for the 2D Heat Exchanger. In green, exists every SAR layer included in the neural operator which does not utilize the optional threshold parameter. The green SAR layers replace either identity mappings or ReLU layers. Components in dark purple represent SAR with the optional threshold parameters τ included for improved spiking control since they follow a normalization layer and non-linear computational layers. These threshold layers replace an identity mapping (spatial) and GeLU (spectral). SAR-GNO takes in boundary input with the input embedding mapping M and combines it with the geometry coordinates to produce an input for the latent projection mapping P . Subsequent spectral-spatial blocks (10 layers total) provide global and local analysis that is combined through a collaboration layer f . Final a downlift layer Q provides the final full-field reconstructed multi-physics output.
Figure 4: Synthetic Distillation We show the synthetic distillation framework utilized for neuromorphic virtual sensing. A graph-based VIRSO model, not native for neuromorphic hardware due to high connectivity and difficult integration, generates synthetic Heat Exchanger examples by randomly sampling input parameters for the Heat Exchanger dataset. These synthetic examples are compared against SAR-NOMAD’s predictions, a more neuromorphic friendly model, allowing for improved L2 error performance in SAR-NOMAD, ideally keeping efficiency the same.
Figure 5: 2D Heat Exchanger (a) We define the 2D Heat Exchanger geometry which is a cross-sectional slice of a 3D grid at axial position x=0.789 with a dimpled wall surface and tape insert utilized for better heat transfer, creating large vortex flow behavior and an irregular geometry with reduced symmetry (b) The boundary input utilized for our 2D Heat Exchanger. It includes a wall heat profile defined along the axial direction of the original 3D geometry and two inlet quantities: temperature and axial velocity/speed. (c) The 50th percentile test example for the SAR-NOMAD model with the Hoyer loss term and γ=0.005 . We see that the main physics and vortex behavior is closely captured with mostly stochastic absolute error behavior and only small absolute error hotspots located around the wall surface.
Figure 6: Spiking Activity Results We depict the probability/frequency of different feature dimensions spiking for the SAR-NOMAD model with the Hoyer loss and γ=0.005 . For the branch layers, the sparsity regularization forces only a small subset, as a low as one dimension, to fire and communication information, indicating that the model is forced to collapse it feature dimensions in order to provide low spiking. The trunk and combined subnetworks have a lot more activity across the entire dimension.
Figure 7: 2D Lid-Driven Cavity (a) We define the boundary conditions for the Lid-Driven Cavity geometry which shows zero velocity on all faces except for a temporal signal at the top surface. (b) We display example time-varying velocity profiles for the top surface which is defined for 90 time steps. (c) The 50th percentile test example for the SAR-NOMAD model with the Hoyer loss term and γ=0.05 . We see that the main vortex and boundary behavior is closely captured as well as the corner eddies with mostly stochastic absolute error behavior. We do see small hotspots of absolute error located at the top surface and corners indicating some difficulty in capturing surface physics with a regular discretized grid.
Predicting full-field physics through the real-time virtual sensing of engineering systems can enhance limited physical sensors but often requires sparse-to-dense reconstruction, complex multiphysics, and highly irregular geometries as well as strict latency and energy constraints for edge-deployability. Neural operators have been presented as a potential candidate for such applications but few architectures exist that explicitly address power consumption. Spiking neuron integration can provide a potential solution when integrated on neuromorphic hardware but the current existing neuron models result in severe performance degradation towards regression-based virtual sensing. To address the performance concerns and edge-constraints, we present the Variable Spiking Graph Neural Operator (VS-GNO) which integrates a sophisticated spectral-spatial convolutional analysis and a previously developed Variable Spiking Neuron (VSN) and energy-error balance loss function. With a non-spiking L2 error baseline of 0.4%, VS-GNO can provide a reconstruction error of 0.71% with 15% average spiking in its spectral-only form and 1.04% with 24.5% spiking in its entire form. These results position VS-GNO as a promising step towards energy-efficient, edge-deployable neural operators for real-time sparse-to-dense virtual sensing in complex, highly irregular engineering environments.
William Howes, Farid Ahmed, Kazuma Kobayashi +2
Grainger College of Engineering, Nuclear, Plasma & Radiological Engineering Department, University of Illinois Urbana-Champaign, Urbana, IL, USA · Department of Applied Mechanics, Indian Institute of Technology Delhi, New Delhi, India · Yardi School of Artificial Intelligence, Indian Institute of Technology Delhi +1
Spiking neural operators are appealing for neuromorphic edge computing because event-driven substrates can, in principle, translate sparse activity into lower latency and energy. Whether that advantage survives deployment on commodity edge-GPU software stacks, however, remains unclear. We study this question on a Jetson Orin Nano 8 GB using five pretrained variable-spiking wavelet neural operator (VS-WNO) checkpoints and five matched dense wavelet neural operator (WNO) checkpoints on the Darcy rectangular benchmark. On a reference-aligned path, VS-WNO exhibits substantial algorithmic sparsity, with mean spike rates decreasing from 54.26% at the first spiking layer to 18.15% at the fourth. On a deployment-style request path, however, this sparsity does not reduce deployed cost: VS-WNO reaches 59.6 ms latency and 228.0 mJ dynamic energy per inference, whereas dense WNO reaches 53.2 ms and 180.7 mJ, while also achieving slightly lower reference-path error (1.77% versus 1.81%). Nsight Systems indicates that the request path remains launch-dominated and dense rather than sparsity-aware: for VS-WNO, cudaLaunchKernel accounts for 81.6% of CUDA API time within the latency window, and dense convolution kernels account for 53.8% of GPU kernel time; dense WNO shows the same pattern. On this Jetson-class GPU stack, spike sparsity is measurable but does not reduce deployed cost because the runtime does not suppress dense work as spike activity decreases.
Jason Yoo, Shailesh Garg, Souvik Chakraborty +1
University of Illinois Urbana-Champaign Department of Nuclear, Plasma & Radiological Engineering Urbana, IL, USA · Indian Institute of Technology Delhi Department of Applied Mechanics New Delhi, India · National Center for Supercomputing Applications Urbana, IL, USA
We propose EdgeSpike, a co-designed spiking neural network (SNN) framework for autonomous low-power sensing in edge Internet of Things (IoT) architectures. EdgeSpike unifies (i) a hybrid surrogate-gradient and direct-encoding training pipeline, (ii) a hardware-aware neural architecture search (NAS) bounded by per-inference energy and memory budgets, (iii) an event-driven runtime targeting Intel Loihi 2, SpiNNaker 2, and commodity ARM Cortex-M microcontrollers with custom spike-sparse SIMD kernels, and (iv) a lightweight local plasticity rule enabling continual on-device adaptation without backpropagation. The framework is evaluated across five sensing tasks (keyword spotting, vibration-based machine fault detection, surface electromyography gesture recognition, 77 GHz radar human-activity classification, and structural-health acoustic-emission monitoring) on three hardware targets. EdgeSpike achieves a mean classification accuracy of 91.4%, within 1.2 percentage points (pp) of strong INT8 convolutional neural network (CNN) baselines (mean 92.6%), while reducing energy per inference by 18x to 47x on neuromorphic hardware (mean 31x) and by 4.6x to 7.9x on Cortex-M (mean 6.1x). End-to-end latency remains at or below 9.4 ms across all 15 task-hardware configurations. A seven-month, 64-node wireless field deployment confirms a 6.3x extension in projected battery lifetime (from 312 to 1978 days at 2 Wh per node) and bounded accuracy degradation under seasonal drift (0.7 pp with on-device adaptation versus 2.1 pp without). Hardware-aware NAS evaluates 8400 candidates and yields a 12-point Pareto front. EdgeSpike will be released as open source with reproducible training pipelines, hardware-portable runtimes, and benchmark suites.
Gustav Olaf Yunus Laitinen-Fredriksson Lundstrom-Imanov, Taner Yilmaz
Department of Economics, Stockholm University, Universitetsvägen 10 A, SE-106 91 Stockholm, Sweden · Department of Computer Engineering, Afyon Kocatepe University, 03200 Afyonkarahisar, Türkiye