Hardware accelerators based on physical dynamical systems offer an attractive route toward energy-efficient reinforcement learning applications. However, their scalability is challenging because it requires many statistically independent entropy sources. Here, we introduce a quasi-analog decision-making architecture based on asynchronous Boolean networks (or lattices) implemented on a clockless reconfigurable chip. Each node in the network consists of a single logic element that acts as an autonomous entropy source. This architecture gives rise to distributed Boolean chaos, in which a spatially coupled network generates parallel streams of chaotic Boolean transitions with very low statistical dependence between nodes. We experimentally demonstrate parallel decision-making on a 1024-armed bandit problem, which is beyond the scale of previous hardware implementations, while significantly improving power-law scaling performance. Separately, we scale the proposed entropy source to 5120 parallel channels, yielding an aggregate sample generation rate of 2.14 TS/s. Our solution is implemented on a commercial reconfigurable CMOS chip and offers high integration density and ease of programmability. Our results pave the way for using distributed Boolean chaos as a valuable hardware substrate for large-scale reinforcement learning and for the development of fully integrated, high-throughput decision-making accelerators.
Figures & tables
Figure 1 : Boolean lattice topology and experimental setup. The adopted topology is a periodic hexagonal lattice, constructed from bidirectionally coupled XOR gates (blue), with a single XNOR gate (red) introduced to drive instability. The lattice is represented on a torus, highlighting the periodic boundary conditions of the lattice. We also represent the repeating hexagonal pattern on the flattened graph. The lattice is implemented on an FPGA, with individual nodes measured using a high-speed oscilloscope after a 10 dB attenuation stage. Nodes to be sampled are selected via control signals from a host computer through a UART interface, and oscilloscope measurements are transferred to the host computer for offline processing.
Figure 2 : Properties of the measured chaotic signals. (a) Measured voltage as a function of time over a 100 ns window. (b) Histogram of pulse lengths ΔT for logic-high (red), logic-low (blue), and combined (gray) states, measured over a 40 μs window. (c) Same as (b), for log2(ΔT) . (d) Pairwise normalized mutual information (NMI) computed from experimentally measured dwell-time distributions of Boolean transitions across the 32-node hexagonal lattice with periodic boundary conditions. Entropies are estimated from empirical probability distributions obtained using histogram binning with 32 bins.
Figure 3 : Solving the MAB problem using Boolean chaos. (a) Decision-making scheme based on parallel chaotic entropy sources. (b) Biased signal produced by the TOW algorithm for a 64-armed bandit problem. The machine with the highest value is selected. (c) Machine selection over successive plays for one realization of the 64-armed bandit problem, alongside relative reward probabilities assigned to each machine (left), with the optimal machine highlighted in red. (d) Unbiased signals from the Boolean chaotic entropy source at play 1000. (e) Signal from (d) after bias adjustment by the TOW algorithm, at the same play.
We propose a scalable neuromorphic architecture based on spiking dynamics emerging from the autonomous time-continuous evolution of clockless (asynchronous) digital circuits. Implemented on commercially available field-programmable gate arrays (FPGAs), our system implements networks of interacting Boolean spiking neurons with configurable excitatory and inhibitory synaptic weights. A complete processing pipeline enables efficient handling of spike-encoded data for solving machine-learning tasks. We demonstrate competitive performance for an audio classification task with spike-based encoding and high-speed processing. Power consumption is significantly lower than traditional digital implementations; this makes our approach an efficient alternative that bridges the gap to dedicated analog neuromorphic systems without the need for specialized hardware design. More generally, our approach establishes clockless digital hardware as a viable platform for neuromorphic computing. It paves the way for reconfigurable chips to be turned into energy-efficient quasi-analog neuromorphic processors.
Eric Oliveira Gomes, Damien Rontani
LMOPS UR4423 Laboratory, CentraleSup´elec and Universit´e de Lorraine, Metz F-57070, France
Pseudo-random number generation often requires trade-offs among quality, power consumption, and bandwidth to produce unpredictable sequences of numbers. The brain, on the other hand, efficiently generates unpredictable output complex network dynamics occurring in a high-dimensional state. This state, which is hypothesized to be chaotic, relies on the balance between excitation and inhibition. Here, we investigated if computational models of these chaotic balanced states can be harnessed for Neuromorphic Pseudo-Random Number Generators (NPRNGs) in low power hardware. We successfully constructed a balanced spiking neural network model consisting of leaky-integrate-and-fire neurons that could be readily implemented in low power FPGAs and used as a NPRNG. The prototyped NPRNG consumed 3.24 mW during operation and produced pseudo-random numbers at 120kbps. In both hardware and software instantiations, NPRNGs produce high-quality random numbers as validated by standard metrics for testing RNG quality.
Jafar Shamsi, Navid Akbari, Sonia Sennik +2
Hotchkiss Brain Institute, University of Calgary, Calgary, Canada · Biomedical Engineering, University of Calgary, Calgary, Canada · Electrical and Software Engineering, University of Calgary, Calgary, Canada +2
Chaotic dynamical systems pose a fundamental challenge for Reinforcement Learning (RL): exponential sensitivity to initial conditions induces high-variance bootstrap targets and poorly conditioned gradient updates. Chaotic dynamics arise across scientific and engineering domains, from fluid flows and climate systems to multi-agent systems, where reliable learning is highly desirable. Standard RL methods optimise expected returns through scalar value functions, implicitly averaging over diverging trajectories and entangling trajectory level instability with the learning objective. We show that under mild statistical stability assumptions, the return distribution evolves more regularly than individual trajectories when measured under the 1-Wasserstein metric, yielding a smoother distributional Bellman objective. By aligning optimisation with this measure level structure, distributional RL provides better conditioned learning. We offer a principled explanation for the advantages of distributional methods in chaotic systems and the geometries of RL objectives under chaos.
James Rudd-Jones, Mirco Musolesi, María Pérez-Ortiz
Centre for Artificial Intelligence Department of Computer Science University College London London, UK · Department of Computer Science and Engineering, University of Bologna Bologna, Italy