Lock-in EP: An In-Situ Training Algorithm for Oscillatory Hardware
Authors: Sowjanya Tammali, Wilkie Olin-Ammentorp
Organizations: Computer Science, Missouri University of Science and Technology, 500 West 15th Street, Rolla, 65409, MO, USA. · Mathematics & Computer Science Division, Argonne National Laboratory, 9700 S. Cass Avenue, Lemont, 60439, IL, USA.
Analog hardware platforms offer the potential to reduce energy consumption over digital architectures, but in order to succeed, large-scale analog systems must also be able to operate with or recover from the variability of their components. Towards this goal, we derive and demonstrate the lock-in equilibrium propagation (LIEP) training method. LIEP provides local gradient information for each component in an oscillatory network without separate forward and backward sweeps, potentially allowing for in-situ learning capabilities on analog oscillatory hardware platforms. We demonstrate that LIEP can be used both for ab-initio training as well as recovering performance when pre-trained parameters are perturbed. We show that LIEP can be formulated as a three-factor update rule, and suggest that although the method is currently only validated on shallow networks, alternate architectures may allow it to extend to deep and large-scale networks addressing complex tasks.
Figures & tables
Figure 1: Layers of neurons with states zℓ are connected via symmetric weights (left). The energy of each layer is defined as the non-alignment in phases between that layer’s current state and its input (Eq. 5 ). For intermediate layers (e.g. z1 ), descending this derived energy function can be accomplished by following the “drive” function (Eq. 10 ) which is dependent on both the previous layer ( zℓ−1 ) and next layer ( zℓ+1 ). The phase term of the state θj aligns with the phase of the drive term θDj , decreasing the layer’s energy over time. The desired outputs y can influence the energy at each layer by contributing an output layer drive from a “virtual” layer ( L+1 ), normalized by the output dimensionality d and modulated by β .
Figure 2: In standard neural networks, the state for a layer is reached by multiplying inputs by weights and applying a nonlinear “squashing” activation function ρ (left). In EP networks, fixed states are reached by following the drive function D at each layer (Eq. 10 ). In LIEP networks, each layer’s oscillatory state zℓ lies on the unit torus in the complex plane, where each point corresponds to set of phase values. To meet this prerequisite, the projection function u maps the drive ( D1 , D2 , D3 ) onto the torus (center). By utilizing a linear update of scaled drive steps (Eq. 12 ), the layer’s state converges to a fixed point on the unit torus over time (right).
Figure 3: Utilizing the network’s energy function, a drive function combined with an appropriate update step allows the network to reach stable, fixed points given an input x , desired output y , and cost weighting term β . When β=0 , the network’s “free” fixed point z∘ is generated. Varying β influences the network’s energy, producing different fixed points when y does not align with the network’s output. Varying β from 0 to a positive value ε is simple (a), but produces a biased estimate of the local gradient around the free fixed point. Adding a second sampling point at −ε reduces the error of this estimate (b). Alternatively, we may modulate β through time following a square-wave schedule to sweep it through the necessary values (c). We may further simplify the probe by replacing the square wave with a simple sinusoid which sweeps β through time (d). This sinusoidal schedule replaces discrete “settling” stages, instead injecting the probe as a time-varying signal with a single harmonic component and allowing the local gradient at the free fixed point to be demodulated.
Figure 4: Applying a time-varying cosine probe β(t) continuously modulates the contribution of the cost function to the network’s energy function (a). As a result, the Hebbian energy term of each layer Hℓ(t) varies through time as well (b). Removing the “DC” energy term of the network at its free fixed point Hℓ∘ and multiplying by the probe to “lock-in,” the resulting signal (c) can be averaged over multiple periods of the probe to estimate the gradient loss function w.r.t. network weights (Eq. 20 ) at each network layer (d). By sampling full periods, the explicit subtraction of the DC signal may also be dropped, as it averages to zero. Inspecting the spectrum of the energy Hℓ , the “DC” free state sits at zero, with the cosine probe creating a clear sideband at its fundamental frequency ωp and higher-order harmonics (e).
Figure 5: We recapitulate the stages of our derivation, and highlight the relevant formulas and quantities passed from each stage to the next. By defining an energy function on the network Φ , we create an associated drive function D which provides “steps” which can lower the energy of the network with inputs x , target outputs y , and weights W (a). The per-layer drive Dℓ is projected onto the unit torus in the complex plane, and a linear update step applies the drive to descend the energy function and “land” fixed network points z∘ on this torus (b). The contribution of the target outputs y to the cost function, β , is varied following a cosine schedule with amplitude ε and frequency ωp (c). The resulting time-varying per-layer Hebbian energy, Hℓ(t) can then be demodulated to provide an estimate of the loss function gradient w.r.t. network weights at each layer (d). These gradients can then be applied to step the network through an outer loop which reduces the loss function over time.
Figure 6: LIEP is controlled by 3 major parameters: the frequency ( ωp ) and amplitude ( ε ) of the cosine probe, and the number of integration periods ( T ) over which the lock-in signal is integrated. The quality of the gradients measured by their cosine similarity to BP-derived values is plotted across a grid where ωp and ε are swept, establishing the existence of an “adiabatic” basin. In this regime, the slow probe frequency and moderate probe amplitude allow the network to settle to steady states and produce high-quality gradient information (a). If the probe becomes too fast, approaching the rate at which the network descends to fixed points ( κ ), the energy dynamics of the system can no longer produce this information, and the ability of the networks to train successfully collapses, showing high failure rates (b) and non-convergence of the loss function (d). Networks trained ab-initio with LIEP in the adiabatic region approach performance levels of standard BP on the same network architectures (c).
Experiment
Fig.
ε
ωp
ncyc
Twarm
Tfree
dt
T
Train
Adiabatic grid
6 a,b
swept
swept
swept
1
200
0.5
varies
—
Basin training
6 c,d
0.003–0.03
0.01–0.10
4
1
200
0.5
varies
60k
Impairment sweep
7 a–d
0.03
0.02
4
2
200
0.5
2512
60k
Three-factor identity
8
0.05
0.05
4
2
100
0.1
5028
—
Table 1: Lock-in probe settings behind each figure. T is the demodulation window in integrator steps, ncyc periods of the probe; a further Twarm periods are discarded before accumulation begins. All runs use 784→256→64 at initial weight scale 0.4 . With κ=0.080 , the ωp=0.02 used for the impairment studies is ωp/κ=0.25 — point B of Fig. 6 , well inside the adiabatic basin and about a factor of two below the fitted boundary. Two settings are not shared across Fig. 7 : its panel (b) uses two warm-up periods and the full 60 k training set, while the paired comparison in (c)–(d) uses one and 10 k. The paired design is internally controlled, but accuracies are not directly comparable between the two.
Figure 7: Rather than competing directly with standard backpropagation (BP) for large-scale training of digital networks, LIEP may be best applied as an analog-native training algorithm for correcting device-level distortions as pretrained networks are deployed onto devices. Many analog devices, such as resistive memories, have limited precision which causes the desired weights to be mapped to a variety of states. This mapping may follow a symmetric Gaussian curve (a), a log-normal scaling (b), or map weights to entirely broken “open” devices (c) or “open/shorted” devices (d). Each of these cases is applied at varying levels of severity, and BP on the entire network or only the readout layer is applied as a control correction. These methods are compared against the experimental LIEP. For many levels of disruption, LIEP can provide recovery approaching that of full-network BP. All accuracies plotted are the median of 8 samples.
Figure 8: LIEP may be reformulated as a local, 3-factor learning rule. We begin by revisiting the final, matrix form of the loss-reducing weight update step (a). The quantities required for this update can be separated into locally-available energy information at each synapse (b) which measures the phase alignment of the pre- and post-synaptic neurons. This alignment is compared to the current state of the global probe (c), which allows the demodulated weight update to be accumulated and integrated into an eligibility trace (d) that can be used to update each weight locally.
Worst
Relative error
Tensor
Components
1−cos
median
max
W1
200 704
4×10−6
1.2×10−3
2.8×10−3
W2
16 384
3×10−7
3.8×10−4
5.4×10−4
Table 2: The numerical results of the three-factor formulation of LIEP are compared against the original, network-level implementation. The same shallow MLP network used previously (Sec. 3.1 ) is optimized using the original LIEP formulation (Sec. 2.2 ) and its 3-factor version derived here. For all components in the network, all match to relative errors less than several parts per thousand.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Nudge β(t)
component at +ωp
component at −ωp
what is extracted
±ε (static)
±ε(a+b) , at DC
a+b , from two or three settles
εeiωpt (complex)
εa
εb
a only, if demodulated at +ωp
εcos(ωpt) (real)
2ε(a+b)
2ε(a+b)
a+b , from one settle
Appendix
Table A.1: Response of the equilibrium to three nudging protocols, to first order in the probe amplitude.
Quantity
Value
Quantity
Value
ρ1=∥W1∥2
6.02
ϱ(G) from min drive
359
ρ2=∥W2∥2
2.10
from 1% quantile
4.90
μ1 (min drive)
0.017
from 10% quantile
2.13
μ2 (min drive)
0.0020
from 25% quantile
1.53
median drive, layer 1
4.21
from median drive
1.09
median drive, layer 2
0.877
required
<1
Appendix
Table B.1: Measured on the trained 784→256→64 network, 500 test samples. ϱ(G) is reported from the drive minimum ( B.13 ) and from four quantiles of the drive distribution; the median row is the most generous reading any refinement of ( B.17 ) could achieve.
Architecture
L
ρL/ρ1
ϱ(G)
out dev
BP
EP
drop
bound
784→256→64
2
0.350
1.095
0.285
0.882
0.870
+0.012
0.162
784→256→256→64
3
0.308
1.838
0.774
0.880
0.828
+0.052
0.782
784→256→256→256→64
4
0.277
2.118
0.975
0.876
0.742
+0.134
0.876
784→64→64
2
0.576
1.331
0.985
0.868
0.818
+0.050
0.846
784→128→64
2
0.417
1.186
0.378
0.862
0.864
−0.002
0.278
784→512→64
2
0.305
1.078
0.262
0.878
0.872
+0.006
0.136
Appendix
Table B.2: Depth and width sweep, seed 42 , 500 test images, five epochs of Adam per configuration. “out dev” is the mean relative deviation δ^L at the output layer, against a trivial maximum of 2 ; “bound” is Proposition B.7 evaluated per sample. Every row satisfies bound ≥ drop.
Department of Electrical and Computer Engineering, Cornell University, New York, NY · IBM T. J. Watson Research Center, Yorktown Heights, NY · Rensselaer Polytechnic Institute, Troy, NY
School of Artificial Intelligence, Beijing Normal University, Beijing 100875, China · Laboratory for Advanced Computing and Intelligence Engineering, Information Engineering University, Zhengzhou 450001, China · School of Physical Science and Technology, Beijing University of Posts and Telecommunications, Beijing 100876, China