Organizations: Hotchkiss Brain Institute, University of Calgary, Calgary, Canada · Biomedical Engineering, University of Calgary, Calgary, Canada · Electrical and Software Engineering, University of Calgary, Calgary, Canada · Creative Destruction Lab · ACHRI, University of Calgary, Calgary, Canada
Pseudo-random number generation often requires trade-offs among quality, power consumption, and bandwidth to produce unpredictable sequences of numbers. The brain, on the other hand, efficiently generates unpredictable output complex network dynamics occurring in a high-dimensional state. This state, which is hypothesized to be chaotic, relies on the balance between excitation and inhibition. Here, we investigated if computational models of these chaotic balanced states can be harnessed for Neuromorphic Pseudo-Random Number Generators (NPRNGs) in low power hardware. We successfully constructed a balanced spiking neural network model consisting of leaky-integrate-and-fire neurons that could be readily implemented in low power FPGAs and used as a NPRNG. The prototyped NPRNG consumed 3.24 mW during operation and produced pseudo-random numbers at 120kbps. In both hardware and software instantiations, NPRNGs produce high-quality random numbers as validated by standard metrics for testing RNG quality.
Figures & tables
Figure 1: Quantizing-Balanced Spiking Neural Networks for Pseudo-Random Number Generation. (A) Schematic of a spiking neural network with random, but balanced connectivity. Neuron i forms excitatory ( +Ng ) and inhibitory synaptic ( −Ng ) connections. (B) The voltage traces (red) and postsynaptic filters (black) for 5 neurons in a simulated balanced network connected as in (A) for 1 second of simulation time. (C) A spike raster plot for 1 second of simulation for all N=256 neurons. (D) The inter-spike-interval (ISI) histogram measured for all neurons in a 100 second simulation of a network with N=256 neurons (from (B)-(C)). The lack of density at 0 is caused by the finite time required for a neuron to generate a spike by integrating a current (the relative refractory period). The distribution of the coefficient of variation (CV) for each of the N=256 is shown (inset). A CV of 1 indicates spiking behaviour similar to that of a Poisson process. (E) The irregular dynamics exhibited by balanced networks are chaotic: if a single spike fails to produce a post-synaptic effect, the subsequent constellation of spikes produced by the network is altered. The non-deleted voltage trace is shown in red, while a spike deletion is performed at 500 ms, for a single neuron (blue). (F) The irregular dynamics exhibited by balanced networks are sensitive to single weight flips. For a pair of parameter-matched, parallel simulations, a single weight ( w12 ) was flipped in sign (excitatory to inhibitory). The voltage trace for the original weight matrix is shown in red, whereas the grey traces show activity in the network before and after the weight flip (vertical line). (G) In a quantized version of the spiking neural network, all network parameters are set to powers of 2 (binary) to enable a simplified implementation of all network operations in hardware. (H) A 1-second simulation of a quantized, balanced neural network, with the voltage traces for 5 neurons was plotted. (I) A 1-second raster plot of a quantized, balanced neural network, for all N=256 neurons. (J) The ISI distribution, and CV distribution (inset) for all of the spikes in a quantized, balanced neural network.
Figure 2: Using Balanced Spiking Neural Networks to Generate Random Bit Strings. (A) Many conventional pseudo-random number generators (PRNGs) use an input seed s and a recursive/iterative operation f(x) to generate a bit string. (B) A neuromorphic pseudorandom number generator (NPRNG) operates differently from a conventional PRNG. Due to hardware limitations, the neural network requires an input signal to initialize the system onto a chaotic trajectory, and generate spikes/spike times pseudorandomly. The NPRNG produces a spike raster plot, which is then transformed into a bit stream with some mapping function, T(s) . (C) The NIST-SP800-22 package of statistical tests is used to test the null hypothesis that the bit string is generated from a series of independent Bernoulli random variables with bk=1 with probability p=1/2 , and 0 otherwise (see Methods for a detailed discussion). (D) (Top) A 5 Hz sinusoid is used as the input signal for a quantized network of N=256 neurons. The signal is shut off after 1 second. (Bottom) A randomly generated sequence of square wave pulses is used as a second input signal for the same network as in (A). (E) NIST testing of a a naive transform, b=T(s) for converting the spike raster plot into a bit stream. For a network of N=2k neurons in base k , each spike index is represented in k -bit binary. NIST testing reveals this as a poor way of generating pseudo-random bits. (F) NIST testing of NPRNG bit streams with dynamic look-up tables for a single weight matrix with different inputs (orange) and different weight matrices (blue). (G) The coefficient of variation of the inter-spike-interval distribution, as a function of the coupling gq (x-axis) and the decay time, τd (y-axis) of the synapses. The bright region in the top-corner corresponds to the “rate-chaos" regime for the SNN. (H) The average number of NIST tests passed as a function of the coupling gq (x-axis) and the decay time, τd (y-axis) of the synapses. (I) Comparing quantized and non-quantized SNN bitstreams to bitstreams generated by other PRNGs.
Figure 3: NPRNGs implemented in hardware with low Power Field Programmable Gate Arrays (FPGAs). (A) Microarchitecture of the NPRNG. (Left) Hardware microarchitecture of the NPRNG. The main microarchitecture (Left) comprises a data path and a control unit. The data path unit includes a memory, multiply-accumulate (MAC) block, synaptic and neuron dynamics, function block F[n] , and the look-up table L[y] . The synaptic and neuron dynamics are implemented using pipeline design. Two, five, and five pipeline stages are used for h[n] , r[n] , and v[n] , respectively. Each block includes a computational block and a RAM. The synapse and neuron blocks perform the computations and the results are stored in the RAM for the next time step. (B) Photograph of the 1cm × 1cm printed circuit board (PCB) with FPGA, two switching regulators, and one diode. (C) Current and power consumption from simulations and tests. The higher power consumption on the fabricated device is due to the extra components and regulator inefficiencies. (D) A Monte-Carlo estimate of π using the hardware-based NPRNG. A sequence of 2.2 million random numbers from the uniform distribution on [0,1] were generated using the NPRNG. These created 1.1 million ordered pairs, uniformly distributed on [0,1]2 (the unit square). The proportion of points that lie below the red line is approximately π/4 . The final Monte-Carlo estimate of π using the NPRNG was 3.1413.
PRNG
Area Metrics (LUT*, FF, DSP)
Frequency (MHz)
Power (mW)
Normalized Power (mW/MHz)
Energy per bit (nJ/bit)
Throughput (bit/clock)
Targeted Hardware
NIST Test
LCG-Based [ 50 ]
(440, 128, 0)
282.64
36.84
0.13
0.13
1
Xilinx Vertix-7
passed
BBS-based [ 50 ]
(539, 181, 0)
161.41
24.53
0.15
10.48
1/69
Xilinx Vertix-7
passed
Chaos-based [ 14 ]
(242, 65, 8)
61.94
114.00
1.87
-
-
Xilinx Artix-7
passed
cNN-based [ 21 ]
(12215, 3117, 40)
0.867 (effective)
-
-
-
8
Altera Cyclone V SoC
passed
NPRNG
(4544, 3441, 0)
6.00
3.26
0.54
27.16
≈ 1/50
Lattice Semiconductor iCE40UP5K
passed
Table 1: Comparison of FPGA implementation of PRNG.
Parameter
Value
Input length (length of a bit stream)
1,000,000
Number of bitstreams
100
Block Frequency Test - block length (M)
128
Non Overlapping Template Test - block length (m)
9
Overlapping Template Test - block length (m)
9
Approximate Entropy Test - block length (m)
10
Table 2: Parameters used for all NIST experiments
Complexity
Memory
CONVERT
MUX
Adder-Tree
h[n]
r[n]
v[n]
y[n]
L[y]
Overall (Max)
Space
O(N2)
O(N)
O(N)
O(log(N))
O(N)
O(N)
O(N)
O(1)
O(N)
O(N2)
Time
O(N)
O(N)
O(N)
O(N)
O(N)
O(N)
O(N)
O(N)
O(N)
O(N)
Table 3: The space and time complexity of the blocks in the microarchitecture.
N
256
256
τm
10 ms
2−7
vreset
-65 mV
0
vpeak
-40 mV
24−1
τD
20 ms
2−6
τR
2 ms
2−9
g
0.1
2−5
Table 4: Table of parameters for spiking neural network simulations.
Supplementary Figure S1 : The behavior of the non-quantized SNN with binarized weights for varying connection strengths (A) The voltage traces for 5 neurons from four networks with increasing g for a 3 second period. Note that all 4 networks have the same weight matrix, only scaled by the different g value ( wij=±Ng ). As g is increased, the neurons fire bursts of varying size/duration. Each network contains N=256 neurons. (B) The distribution of inter-spike-intervals for the 4 simulated networks from (A). For lower values of g , the distributions resemble a Poisson process with a refractory period. Note that the log-ISI count is used on the y-axis. (C) The distribution of the coefficient of variation for the networks in (A)-(B). The mean and standard deviation of the CVs are shown in the title. For more strongly coupled networks g=0.15,g=0.2 , the mean CV is greater than 1, indicating the burstiness of the neurons in these coupling regimes.
Supplementary Figure S2 : The behavior of a quantized SNN with binarized weights for varying connection strengths. (A) The voltage traces for 5 neurons from four networks with increasing quantized connection strength gq=2−n for a 3 second period. Note that all 4 networks have the same weight matrix, only scaled by the different gq value. As gq is increased, the neurons fire bursts of varying size/duration. Each network contains N=256 neurons. All operations in the network are quantized. (B) The distribution of inter-spike-intervals for the 4 simulated networks from (A). For lower values of gq , the distributions resemble a Poisson process with a refractory period. Note that the log-ISI count is used on the y-axis. (C) The distribution of the coefficient of variation for the networks in (A)-(B). The mean and standard deviation of the CVs are shown in the title. For more strongly coupled networks gq=2−8,gq=2−7 , the mean CV is greater than 1, indicating the burstiness of the neurons in these coupling regimes.
Supplementary Figure S3 : Synchronizing NPRNGs with input currents. (A) The voltage traces for 5 neurons, of the same balanced, and quantized network of N=256 neurons. Voltages for each neuron are initialized at different values for both simulations (red, black). An input current is applied to each neuron, multiplied by an input weight (Materials and Methods). The input current is non-zero for a short duration of time (250 ms) The input current causes a transient synchronization of the neurons in the network. Eventually, the neuron’s desynchronize. (B) Identical to (A), only the input current is applied for a longer duration of time (750 ms). The neurons maintain synchrony, despite the chaotic dynamics, in perpetuity.
Supplementary Figure S4 : Correlation structure in a Balanced-Quantized Network. (A) The sign of the weights between 5 neurons in a network of N=256 neurons with balanced connectivity, and quantized operations. (B) The autocorrelation (diagonal) and cross-correlations (off-diagonal) between neurons in a balanced, quantized network simulated for T=1000 seconds. The cross-correlations are weak, and scale-like N−1 . The autocorrelation functions have prominent peaks in between a refractory period of non-firing centered at 0.
Supplementary Figure S5 : Average CV versus the average number of NIST tests passed. The mean CV is inversely correlated with the quality of random numbers generated ( ρ=−0.9029 , p≪10−4 , N=128 , ρ=−0.6818 , p≪10−4 , N=256 , ρ=−0.8147 , p≪10−4 , N=128 , ρ=−0.9538 , p≪10−4 , N=512 ). The slops and intercepts for the line of best fits are -0.2576, -0.2684, -0.2256, and 12.0134, 13.9926, and 14.6775, for N=128 , N=256 , and N=512 respectively.
Supplementary Figure S6 : Poorly generated but still balanced excitatory/inhibitory synaptic weights lead to cycling. (A) The weight matrix is generated with 50% excitatory and 50% inhibitory connections. Each row is a circular shift of the previous row, one index to the right. (B) The spike raster plot for a simulated quantized network of neurons coupled with the weights in (A). The network eventually stabilizes on a fixed periodic sequence of spiking. (C) The index of the bit reported in the look-up table. The bits reported in this case also cycle, but at twice the period of the neurons in the network.
Supplementary Figure S7 : The maximum Lyapunov exponent as a function of N , the network size, for increasingly larger N . The coupling parameter was scaled as g=N2−6 . (Left) The proportion of simulations where the spiking eventually ceased, as the non-spiking state is bistable with other attractors. (Right) The numerically computed maximum Lyapunov exponent as a function of N . The red line denotes the 0 point.
Supplementary Figure S8 : Heat map of the average number of NIST tests passed (left) and the coefficient of variations of the resulting spike trains (right) for N=128 (top row), N=256 (middle row), and N=512 (bottom row).
Supplementary Figure S9 : (A) Schematic of inhibition of neurons via bias currents. An externally inhibited neuron (blue) no longer transmit spikes in the network. This is equivalent to setting all of the post-synaptic weights for the inhibited inside the weight matrix to 0. This generates a novel NPRNG from an existing NPRNG. (B) A simulation of the binarized/quantized spiking neural network with 2 neurons inhibited (black) and no inhibitory currents (red). The neurons are inhibited by providing a sub-threshold bias current ( I=0 ). The spike-raster of the two networks is different. (C) Neuronal inhibition allows and chaotic dynamics allows for multiple NPRNGs to be generated from a single network. (D) NIST testing of networks of N=256 neurons with N=3 neurons randomly inhibited for the duration of each simulation. Deleting neurons does not appreciably change PRNG quality.
Stochastic spiking neurons trade exact arithmetic for controlled randomness, lowering area and tolerating input noise, which suits event-driven edge hardware. We present a compact, configurable stochastic leaky integrate-and-fire neuron in standard-cell CMOS on the SkyWater 130 nm process, released openly. A 16-bit configurable-polynomial linear-feedback shift register drives an eight-entry programmable activation table that sets a Bernoulli firing probability, and a saturating 16-bit leaky integrator with a programmable threshold and a refractory period of zero to seven cycles produces the spike train. All parameters are set through a sixteen-register serial interface, and the neuron runs from parallel inputs or entirely from the register file. From a model checked bit-exact against the register-transfer code, the period is 65535 states for a maximal-length polynomial and 63 for the shipped default, the eight-bit comparison value is uniform over the full period, and the per-entry firing probability equals the table value divided by 256. We also characterise a property a system-level model would not expose: the comparator output is serially correlated at short lags, with a negative lobe near lag eight, because the compared byte shifts by one bit each cycle; subsampling every sixteen cycles restores whiteness. Rate-coding sweeps show monotonic control of the output rate by the input weight and the threshold, and the refractory period caps the rate at one spike per refractory-plus-one cycles. The neuron occupies about 10,600 square micrometres at 70 per cent utilisation on a single Tiny Tapeout tile, meets 50 MHz timing with positive margin, and passes eighteen directed cocotb tests at register-transfer and gate level. All results are pre-silicon, from simulation and the open flow. The neuron is an openly released companion to a four-block neuromorphic suite reported separately.
Poornima Kumaresan, Santhosh Sivasubramani
Intrinsic Lab, Centre for Sensors, Instrumentation and Cyber-Physical System Engineering (SeNSE), Indian Institute of Technology Delhi, New Delhi 110016, India
Bayesian Neural Networks (BNNs) offer opportunities for greatly enhancing the trustworthiness of conventional neural networks by monitoring the uncertainties in decision-making. A significant drawback for BNN inference at the extreme edge, however, is the imperative need to incorporate Gaussian Random Number Generators (GRNG) within each neuron. State-of-the-art GRNG algorithms heavily depend on multiple arithmetic operations and the use of extensive look-up tables, posing significant implementation challenges for ultra-low power hardware implementations. To overcome this, this paper presents an innovative binary tree random number generator (TreeGRNG) allowing the use of ultra-low-cost constant comparators instead of arithmetic units. We further enhance the TreeGRNG proposal with a set of hardware-aware optimizations exploiting the Gaussian properties. The optimized TreeGRNG surpasses the State-of-the-Art (SoTA) in terms of distribution accuracy while achieving a 3.7× reduction in energy per sample and boosting the throughput per unit area by 5.8×. Moreover, our TreeGRNG proposal possesses a distinct advantage over the current SoTA in terms of flexibility, as it easily enables designers to adjust the shape of the sampled probability distribution, extending beyond the capabilities of traditional GRNGs, opening the horizon towards future probabilistic AI designs. The TreeGRNG design is available open-source in the link
We propose a scalable neuromorphic architecture based on spiking dynamics emerging from the autonomous time-continuous evolution of clockless (asynchronous) digital circuits. Implemented on commercially available field-programmable gate arrays (FPGAs), our system implements networks of interacting Boolean spiking neurons with configurable excitatory and inhibitory synaptic weights. A complete processing pipeline enables efficient handling of spike-encoded data for solving machine-learning tasks. We demonstrate competitive performance for an audio classification task with spike-based encoding and high-speed processing. Power consumption is significantly lower than traditional digital implementations; this makes our approach an efficient alternative that bridges the gap to dedicated analog neuromorphic systems without the need for specialized hardware design. More generally, our approach establishes clockless digital hardware as a viable platform for neuromorphic computing. It paves the way for reconfigurable chips to be turned into energy-efficient quasi-analog neuromorphic processors.
Eric Oliveira Gomes, Damien Rontani
LMOPS UR4423 Laboratory, CentraleSup´elec and Universit´e de Lorraine, Metz F-57070, France