Organizations: Technion – Israel Institute of Technology, Haifa, Israel · NVIDIA · Institute of Science and Technology Austria (ISTA), Klosterneuburg, Austria
Implicit neural representations (INRs) are shaped by the spectral structure induced by their input encodings and activation functions. Existing methods improve fitting primarily by modifying which frequencies are available to the network, through coordinate encodings or periodic nonlinearities. However, frequency access is not the only bottleneck: signals with localized or spatially varying structure require the network to efficiently compose frequencies into multi-harmonic internal responses. We introduce learnable spectral activations (LSA), which replace fixed neuron-level nonlinearities with a residual truncated Fourier series whose harmonic amplitudes are learned during training. LSA does not expand the asymptotic function class. Instead, it changes the factorization of the representation: linear weights select features while activation coefficients control spectral shaping, and the two are updated by separate gradients. Because the activation output is affine in the coefficients given fixed pre-activations, spectral tuning becomes a more direct subproblem compared to architectures where it is entangled with feature selection. Empirically, this factorization concentrates more target-signal energy in the leading eigenmodes of the neural tangent kernel, consistent with improved optimization behavior. Across audio, image, neural radiance field, and neural acoustic field tasks, LSA also improves reconstruction quality.
Figures & tables
Figure 1 : (a) Operator families and parameter efficiency. Dashed curves show individual atoms and solid curves show the fitted signal. SIREN and FINER use prescribed single-unit responses, whereas LSA synthesizes multi-harmonic waveforms and reaches accurate fits with fewer atoms. (b) Target energy concentration in leading empirical NTK modes. LSA places more target energy in the top modes than SIREN or FINER, indicating better alignment between the induced dictionary and the target.
Figure 2 : SIREN/FINER/LSA comparison of on-grid audio fitting across Fourier-feature bandwidths σ . LSA achieves the strongest performance at moderate and high bandwidths on both speech and musical-note signals, indicating that learned spectral activations improve the internal processing of Fourier-feature harmonics.
Figure 3 : On-grid vs. off-grid reconstruction under increasing Fourier-feature bandwidth. SIREN exhibits a large on-grid/off-grid gap already at σ=0 and severe interpolation artifacts for σ=100 . Our method remains stable under inference of interpolated, unseen grid points.
Figure 4 : Per-image Kodak fitting comparison under fixed configurations. Per-image PSNR is shown for LSA, FINER, SIREN, and STAF under 5,000-update fits.
Figure 5 : Poisson reconstruction from gradient-only supervision. Our method produces more consistent gradient fields and integrates them into a coherent image, while FINER and SIREN exhibit structured residuals and global reconstruction errors.
Figure 6 : Qualitative Kodak reconstructions for illustrative examples kodim21 (left) and kodim13 (right), selected for qualitative clarity. LSA, FINER, and SIREN use the fixed configurations in Figure 4 . We show reconstructions, spatial residuals, and spectral residuals, with a common display scale for each residual type across both images and all methods. Full-dataset comparisons are in Figure 4 and Section C.1 .
Metric
Method
Chair
Drums
Ficus
Hotdog
Lego
Materials
Mic
Ship
PSNR ↑
SIREN [ 2 ]
33.48
24.89
28.22
32.85
29.60
27.13
33.28
22.48
FINER [ 24 ]
33.90
24.90
28.70
33.05
30.04
27.05
33.96
22.47
Ours
34.18
24.79
28.85
33.36
29.39
27.22
33.91
22.60
SSIM ↑
SIREN [ 2 ]
0.973
0.912
0.959
0.960
0.948
0.932
0.979
0.797
FINER [ 24 ]
0.973
0.911
0.958
0.959
0.951
0.928
0.981
0.792
Ours
0.977
0.922
0.967
0.967
0.950
0.936
0.983
0.803
Table 1 : Controlled activation comparison on novel view synthesis, over Blender [ 3 ] scenes.
Figure 7 : Spatial improvement of LSA over the NAF baseline. Each point corresponds to a spatial query in the scene, and color indicates relative improvement. LSA yields broad spatial gains, especially for log-spectral and multi-resolution STFT metrics, across the queried locations.
Table 2 : Audio fitting hyperparameters. FF16 denotes M=16 Gaussian Fourier-feature bands. The selected σ values are those used for the main audio fitting results. Full σ -sweep curves are reported separately.
α/α0
0.10
0.03
0.01
Mean R(α)
0.806
0.942
0.981
Appendix
Table 3 : One-step validation on six Kodak images and three initializations per image. The normalization in Equation 26 tests the local descent prediction.
Figure 8 : Qualitative waveform fits on LibriSpeech and NSynth. Each column compares the target with a reconstruction and its residual on the original audio scale. Learned spectral activations closely match transient and harmonic waveform structure, leaving substantially smaller residuals than FINER and SIREN.
Figure 9 : Full NSynth on-grid vs. off-grid diagnostic across bandwidth settings.
Figure 10 : Convergence curves for all 24 Kodak images under fixed LSA, FINER, and SIREN configurations and matched 5,000-update budgets. LSA separates from SIREN early, while its advantage over FINER emerges progressively. Early-stage behavior varies across images.
Figure 11 : Qualitative novel view synthesis results. We show target views, reconstructions, and absolute RGB residuals for representative scenes. Learned spectral activations reduce structured residuals around fine object boundaries and high-detail regions compared to SIREN and FINER.
Macro-scene aggregation
Micro-position aggregation
Metric
Baseline
LSA
Improvement
Baseline
LSA
Improvement
PSNR from MSE ↑
11.417
11.846
+0.429 / +3.76%
11.550
11.993
+0.443 / +3.84%
MSE ↓
0.07238
0.06572
+0.0067 / +9.20%
0.07
0.063
+0.0068 / +9.71%
Log Spectral L1 ↓
4.053
2.233
+1.820 / +44.90%
3.981
1.957
+2.024 / +50.84%
MRSTFT Log-L1 ↓
4.008
2.165
+1.842 / +45.97%
3.925
1.902
+2.023 / +51.54%
TDOA ↓
122.406
39.727
+82.679 / +67.55%
114.675
31.634
+83.041 / +72.41%
Appendix
Table 4 : Neural acoustic field results. We replace the activation in the NAF model with LSA ( K=8 ) while keeping the architecture and training protocol fixed. Improvements are signed so that positive values indicate better performance.
Figure 12 : Spatial improvement of LSA over the NAF baseline. frl apartment scene 4
Figure 13 : Spatial improvement of LSA over the NAF baseline. Apartment Scene 1
Figure 14 : Spatial improvement of LSA over the NAF baseline. Apartment Scene 2.
Figure 15 : Spatial improvement of LSA over the NAF baseline.Room Scene.
Figure 16 : Spatial improvement of LSA over the NAF baseline. Office Scene
Figure 17 : Sparse context → dense reconstruction under different positional encodings. Top: reconstructions. Bottom: absolute error maps.
Method
Candidates/image
PSNR (dB)
SSIM
PSNR wins
LSA
3
38.308 ± 2.849
0.9568 ± 0.0111
24/24
STAF
56
35.425 ± 2.190
0.9209 ± 0.0119
0/24
Appendix
Table 5 : Per-image Kodak tuning. Values are mean and standard deviation across 24 images. Both methods select their configurations separately for each image.
Method
Parameters
Holdout PSNR
All-24 PSNR
LSA wins
LSA
199,043
38.045
38.110
-
SL 2 A, full
330,243
37.857
37.298
10/18
SL 2 A, parameter matched
199,171
35.910
35.995
18/18
SL 2 A, author protocol
330,243
37.463
37.541
15/18
Appendix
Table 6 : SL 2 A comparison on Kodak. The 18-image holdout is the primary comparison after calibration. The 24-image means also include the six calibration images. Wins refer to LSA relative to each row on the holdout.
Method
Holdout PSNR
LSA gain (dB)
95% interval
LSA wins
LSA
38.045
-
-
-
fJNB
22.985
15.060
[14.595, 15.559]
18/18
WIRE
36.477
1.568
[1.143, 1.970]
17/18
MFN GaborNet
35.086
2.959
[2.764, 3.157]
18/18
Appendix
Table 7 : Additional Kodak holdout results after six-image calibration. Confidence intervals are paired bootstrap intervals for LSA’s mean PSNR gain. Means and paired gains are rounded independently.
Method
Mean PSNR (dB)
SIREN
34.990
FINER
37.317
STAF
35.146
WIRE
36.596
MFN GaborNet
35.151
fJNB
22.948
Appendix
Table 8 : Descriptive Kodak results over all 24 images with globally selected configurations. Calibration images are included. STAF, WIRE, and MFN receive up to 10,000 updates. LSA, SIREN, and FINER use 5,000.
Image
LSA PSNR
STAF PSNR
LSA SSIM
STAF SSIM
kodim01
35.671
32.973
0.9562
0.9240
kodim02
39.563
35.636
0.9529
0.8902
kodim03
42.714
38.166
0.9737
0.9321
kodim04
39.194
35.999
0.9563
0.9086
kodim05
35.157
31.669
0.9605
0.9157
kodim06
38.192
34.668
0.9614
0.9256
Appendix
Table 9 : Per-image STAF comparison after selecting from 56 STAF and three LSA configurations per image.
Input
Terms τ
NSynth
LibriSpeech
Raw coordinates
5
68.79±3.04
68.40±6.29
FF16, σ=20
5
55.11±11.63
50.24±1.38
FF16, σ=100
5
51.83±4.47
51.85±1.28
FF16, σ=100
16
50.60±2.44
51.21±1.43
Appendix
Table 10 : STAF audio configurations. Entries are mean PSNR and standard deviation across clips. The raw-coordinate setting uses first-layer ω0=3000 .
Degree
Rank
Parameters
NSynth
LibriSpeech
512
128
329,473
68.03±10.20
47.42±6.22
512
64
231,169
67.16±10.57
46.35±6.62
256
128
263,937
59.16±16.05
38.84±7.18
256
64
165,633
59.47±15.35
38.37±7.28
Appendix
Table 11 : SL 2 A audio results for the tested degree and rank settings. Values are mean PSNR and standard deviation across clips. All rows use learning rate 10−3 .
Dataset
Baseline
Input
Baseline PSNR
LSA PSNR
Gain
Wins
NSynth
fJNB
Raw
16.45
47.41
30.96
9/10
LibriSpeech
fJNB
Raw
25.23
31.96
6.73
31/40
NSynth
fJNB
FF16, σ=100
55.09
86.72
31.63
10/10
LibriSpeech
fJNB
FF16, σ=100
43.97
90.78
46.81
40/40
NSynth
WIRE
Selected
47.11
86.72
39.61
10/10
LibriSpeech
WIRE
Selected
51.79
90.78
38.99
40/40
Appendix
Table 12 : Additional audio fitting comparisons. Gains and wins refer to LSA relative to the specified baseline on the same clips. fJNB uses the fractional Jacobi construction.
Scene
WIRE (published)
fJNB (adapted)
LSA
Chair
29.31
30.15
34.18
Drums
22.22
18.08
24.79
Ficus
25.91
22.24
28.85
Hotdog
30.11
31.81
33.36
Lego
25.76
22.67
29.39
Materials
25.05
24.73
27.22
Appendix
Table 13 : Additional Blender PSNR results. WIRE is reproduced from Liu et al. [24] . fJNB and LSA use the shared-backbone evaluation described above. Means are across eight scenes.
K
Mean PSNR
Std. PSNR
Median PSNR
Mean SSIM
4
30.44899
5.28322
29.17142
0.81556
8
35.35890
4.63919
34.52907
0.92290
16
38.67377
4.34150
40.33058
0.96070
32
39.50432
4.32379
41.38003
0.96614
64
39.34945
4.14439
40.49803
0.96494
Appendix
Table 14 : Harmonic-count ablation on six Kodak images. Standard deviations and medians are computed across images.
Evaluation
Coefficients
PSNR
SSIM
LSA wins
Gain 95% interval
18 holdout
Learned
38.05±2.63
0.9545
18/18
[17.45,18.87]
18 holdout
Frozen
19.91±1.78
0.4861
24 pooled
Learned
38.11±2.90
0.9557
24/24
[17.56,19.17]
24 pooled
Frozen
19.76±1.95
0.4834
Appendix
Table 15 : Learning versus freezing the near-identity coefficients. PSNR is mean and standard deviation across images. Confidence intervals refer to the paired mean PSNR gain.
Figure 18 : Hidden-representation diagnostics on image fitting. For each method, we visualize layer-wise RMS hidden activity, the two highest-variance signed hidden channels, and the radial spectra of those selected channels. LSA produces more spatially localized and less purely periodic hidden features than SIREN and FINER, supporting the view that learned spectral activations modify the internal basis rather than only the final reconstruction.
Figure 19 : Learned LSA coefficient spectra across Kodak images. Each panel shows log10(∣ak(ℓ)∣+10−8) for harmonic index k and layer ℓ . The learned operators exhibit layer-dependent harmonic profiles and image-dependent variation, indicating that LSA adapts its internal spectral basis rather than using a fixed spectrum.
Figure 20 : Learned activation coefficient spectra across Kodak images. Each panel shows the magnitude ∣ak(ℓ)∣ of the learned harmonic coefficients for one image and layer. The coefficients exhibit layer-dependent profiles and image-dependent variation, indicating that LSA learns nontrivial spectral operators rather than using a fixed harmonic profile.
Figure 21 : Learned activation waveforms across Kodak images. Each panel plots the resulting activation ϕ(ℓ)(u) for one image and layer. Although the learned functions remain close to the residual identity shape, they contain localized harmonic deviations whose form varies across layers and fitted images.
Depth
Method
Params
Extra act. params
Train time (ms)
Train mem. (GB)
Eval time (ms)
Eval mem. (GB)
4
SIREN FF16
206,337
0
8.26
0.48
3.56
0.16
FINER FF16
206,337
0
14.20
0.85
5.55
0.21
LSA FF16, K=16
206,401
64
97.93
7.53
35.08
1.58
6
SIREN FF16
337,921
0
12.77
0.67
5.46
0.16
FINER FF16
337,921
0
21.65
1.22
8.45
0.21
LSA FF16, K=16
338,017
96
147.84
10.56
52.87
1.58
Appendix
Table 16 : Resource usage under the matched FF16 input regime on full-grid 48k audio fitting. All models use width 256. Training time includes forward, backward, and optimizer update. Evaluation time is forward-only. Memory is peak CUDA allocation.
Depth
K
Train time (ms)
Train mem. (GB)
Eval time (ms)
Eval mem. (GB)
4
8
52.78
3.88
19.24
0.85
16
98.57
7.53
35.33
1.58
32
189.87
14.87
67.26
3.05
64
341.09
29.55
115.23
5.98
6
8
79.13
5.44
28.77
0.85
16
147.84
10.56
52.87
1.58
Appendix
Table 17 : LSA scaling with the number of harmonic terms K under the matched FF16 full-grid 48k setting. All models use width 256. The implementation is the batched PyTorch harmonic evaluation used in Table 16 .
K
Implementation
Train time (ms)
Train mem. (GB)
Eval time (ms)
Eval mem. (GB)
8
Batched
6.63
0.35
2.26
0.09
Accumulate
6.98
0.30
2.35
0.04
16
Batched
11.43
0.66
4.08
0.15
Accumulate
13.13
0.55
4.40
0.04
Appendix
Table 18 : Implementation diagnostic for LSA in a no-FF, small-batch setting ( B=4096 , width 256, depth 4, fp32). The accumulate implementation avoids explicitly materializing the full harmonic tensor, reducing memory at the cost of additional runtime.
Implicit Neural Representations (INRs) have been proven successful in encoding continuous signals through coordinate-based networks, yet facing a spectral dilemma: periodic activations capture fine details but act as all-pass filters that memorise noise, while spatially compact activations regularise effectively but suffer from low-frequency bias. Existing attempts to resolve this trade-off introduce computational overhead or tuning frailty. We propose to model each neuron's activation as the steady-state response of a sinusoidally-forced damped harmonic oscillator, whose amplitude naturally governs the network's spectral selectivity during training. By jointly optimising the oscillator parameters alongside the network weights, our method adapts to the target signal's spectral content without explicit regularisation. Initialised in the stopband, the network exhibits a coarse-to-fine learning curriculum that progressively expands its spectral gate, capturing low-frequency structures first and high-frequency details only when justified by the reconstruction objective. Comprehensive experiments show that our approach consistently achieves state-of-the-art or competitive results against established INRs, while requiring no task-specific tuning of any hyperparameters.
Alex Costanzino, Pierluigi Zama Ramirez, Giuseppe Lisanti +1
CVLab, University of Bologna · Ca’ Foscari University of Venice
We study sinusoidal recurrence as an iterative mechanism for harmonic spectral enrichment in implicit neural representations (INRs). Our analysis reveals that sinusoidal activations induce a harmonic line spectrum, providing a spectral account of how recurrent unrolling enriches the effective spectral support. We realize this principle with a shared sinusoidal block that iteratively refines the latent representation. We empirically validate the resulting spectral behavior against feed-forward INRs, non-sinusoidal recurrent variants, and equilibrium-style sinusoidal models. Complementing this analysis, we evaluate the proposed architecture across image and 3D representation tasks. On RGB image benchmarks, our method achieves higher fidelity than feed-forward baselines with fewer parameters and fewer optimization steps, and it further transfers favorably to super-resolution, NeRF, and SDF tasks.
Hyunmin Cho, Jaejun Yoo, Kyong Hwan Jin
Department of Electrical Engineering, Korea University, Seoul, South Korea · Graduate School of Artificial Intelligence, UNIST, Ulsan, South Korea
Implicit Neural Representations (INRs) parameterized by multilayer perceptrons excel at modeling continuous signals. However, a key challenge persists as INRs fundamentally suffer from spectral bias and information cross-talk. When a single network attempts to capture multi-scale phenomena, high-frequency weight updates destructively interfere with the underlying low-frequency structural approximation. We introduce Scale and Learn INR (ScaLe-INR), a novel multi-branch architecture that resolves these limitations by explicitly matching the signal's frequency spectrum with the optimal operating region of the INR. Drawing upon the Fourier inverse scaling theorem we demonstrate that applying directional coordinate scaling expands a network's representational bandwidth along specific spatial axes. To mathematically enforce functional disentanglement and minimize task-specific information leakage between branches, we propose a Directional Edge Guidance Loss, a spatially-conditioned sparsity prior derived from ground-truth gradients. By constraining the high-frequency branches to act as strict, localized edge-filters, ScaLe-INR eliminates spectral cross-talk, accelerates convergence, and achieves high-fidelity signal reconstruction on complex multi-scale topologies. We evaluate ScaLe-INR across diverse reconstruction and inverse tasks, demonstrating substantial performance gains over existing state-of-the-art (SOTA) methods. The proposed architecture improves upon the nearest baselines by +5.16 dB in image reconstruction and +0.65 dB in image denoising. Furthermore, it achieve an impressive figure of 50.02 dB on audio reconstruction and 0.999 IOU(Intersection Over Union) on 3D reconstruction which beats the all SOTA models.