Neural Processes

Momentum

2 papers in the last four weeks, against 1 the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 46

May 7, 2026cs.LG

From Drops to Grid: Noise-Aware Spatio-Temporal Neural Process for Rainfall Estimation

High-resolution rainfall observations are crucial for weather forecasting, water management, and hazard mitigation. Traditional operational measurements are often biased and low-resolution, limiting their ability to capture local rainfall. Accurate high-resolution rainfall maps require integrating sparse surface observations, yet existing deep learning densification methods are hindered by rainfall's skewed, localized nature, noise, and limited spatio-temporal fusion. We present DropsToGrid, a Neural Process-based method that generates dense rainfall fields by fusing temporal sequences from noisy, irregularly distributed private weather stations with spatial context from radar. Leveraging multi-scale feature extraction, temporal attention, and multi-modal fusion, the model produces stochastic, continuous rainfall estimates and explicitly quantifies uncertainty. Evaluations on real-world datasets demonstrate that DropsToGrid outperforms both operational and deep learning baselines, generating accurate high-resolution rainfall maps with well-calibrated uncertainty, even when only few stations are available and in cross-regional scenarios.
May 5, 2026cs.NE

Interpreting V1 Population Activity via Image-Neural Latent Representation Alignment

Understanding the neural mechanisms underlying visual computation has long been a central challenge in neuroscience. Recent alignment based approaches have improved the accuracy of decoding visual stimuli from brain activity, yet they provide limited insight into the neural computations that give rise to these improvements. To address this gap, we propose Dual-Tower Image-Neural Alignment (DINA), an interpretable contrastive framework for analyzing population level visual computations in primary visual cortex (V1). DINA jointly trains a biologically motivated dual-tower architecture that aligns visual stimuli and corresponding V1 population responses in a shared latent space at the level of intermediate feature maps, enabling both accurate decoding and direct access to interpretable feature maps. Evaluated on large-scale two-photon calcium imaging data from mouse V1, DINA achieves accurate neural-based decoding while revealing that decoding performance is primarily supported by coarse, low-level visual structure, rather than semantic category information or fine-grained details. Further analysis reveals that alignable feature maps emerge from multiple spatially distributed image regions, capturing both shape and texture cues, and are predominantly reconstructed by sparse subsets of strongly responsive neurons and their functional interactions. Together, these results confirm that, beyond enabling accurate decoding, DINA provides a principled framework for probing the computational mechanisms underlying visual processing in V1.
May 5, 2026cs.LG

Learning reveals invisible structure in low-rank RNNs

Learning in neural systems arises from synaptic changes that reshape the representations underlying behavior. While low-rank recurrent neural networks (RNNs) have emerged as a powerful framework for linking connectivity to function, a theoretical understanding of their learning process remains elusive. Here, we extend the low-rank framework from activity to learning by deriving gradient-descent dynamics directly in a reduced overlap space. We formulate a closed-form, low-dimensional system of ODEs that governs learning in this space, exact for linear RNNs and asymptotically exact for nonlinear RNNs in the large-NN Gaussian limit. Central to our analysis is a distinction between two classes of overlaps: loss-visible overlaps, which fully determine network activity, output, and loss, and loss-invisible overlaps, which do not affect function but are required to describe learning. We illustrate the consequences of this decomposition through two phenomena. First, we show that learning can serve as a perturbation that exposes differences in connectivity between functionally equivalent networks. Second, we show that loss-invisible overlaps can act as memory variables that encode training history, and characterize the conditions under which this occurs. Finally, we present several testable predictions for biological learning experiments derived from our theory.
Apr 29, 2026cs.LG

Causal Learning with Neural Assemblies

Can Neural Assemblies -- groups of neurons that fire together and strengthen through co-activation -- learn the direction of causal influence between variables? While established as a computationally general substrate for classification, parsing, and planning, neural assemblies have not yet been shown to internalize causal directionality. We demonstrate that the inherent operations of neural assemblies -- projection, local plasticity control, and sparse winner selection -- are sufficient for directional learning. We introduce DIRECT (DIRectional Edge Coupling/Training), a mechanism that co-activates source and target assemblies under an adaptive gain schedule to internalize directed relations. Unlike backpropagation-based methods, DIRECT relies solely on local plasticity, making the resulting causal claims auditable at the mechanism level. Our findings are verified through a dual-readout validation strategy: (i) synaptic-strength asymmetry, measuring the emergent weight gap between forward and reverse links, and (ii) functional propagation overlap, quantifying the reliability of directional signal flow. Across multiple domains, the framework achieves perfect structural recovery under a supervised, known-structure setting. These results establish neural assemblies as an auditable bridge between biologically plausible dynamics and formal causal models, offering an "explainable by design" framework where causal claims are traceable to specific neural winners and synaptic asymmetries.
Apr 28, 2026physics.data-an

Emergent Self-Attention from Astrocyte-Gated Associative Memory Dynamics

We introduce a Hopfield-type associative memory in which effective connectivity is multiplicatively modulated by astrocytic gains evolving under an entropy-regularized replicator equation. The coupled neuron-astrocyte dynamics admit a Lyapunov function, ensuring global convergence. At fixed points, astrocytic gains implement a softmax-normalized allocation over pattern similarity scores, yielding a mechanistic realization of self-attention as emergent routing on the gain simplex. In regimes of high memory load and interference, the model significantly improves retrieval accuracy relative to classical Hopfield dynamics and recent neuron-astrocyte baselines. These results establish a dynamical systems framework linking glial modulation, competitive resource allocation, and attention-like computation.
Apr 26, 2026cs.AI

ZenBrain: A Neuroscience-Inspired 7-Layer Memory Architecture for Autonomous AI Systems

On LongMemEval-500, ZenBrain matches a long-context oracle's binary-judge accuracy to within 4.5 pp (47.7%47.7\% vs. 52.2%52.2\%; 91.3%91.3\%) at 1/106th1/106^\text{th} of the per-query token cost (App. F.5-F.6, Fig. 2), and wins all 12 head-to-head answer-quality cells (4 systems ×\times 3 LLM judges) against Letta, Mem0, and A-Mem under Bonferroni correction (α=0.05/18α=0.05/18, pmin=6.2×10−31p_\text{min}=6.2\times 10^{-31}, d∈[0.18,0.52]d \in [0.18, 0.52]). ZenBrain is a 7-layer neuroscience-inspired memory architecture. The contribution is architectural integration: 15 validated neuroscience mechanisms unified under a single MemoryCoordinator -- 9 foundational algorithms (Two-Factor Synaptic KG, vmPFC-coupled FSRS, Simulation-Selection sleep, Bayesian confidence, and five more) plus 6 Predictive Memory Architecture components (NeuromodulatorEngine, ReconsolidationEngine, TripleCopyMemory, PriorityMap, StabilityProtector, MetacognitiveMonitor). No prior system integrates more than two. Stress ablation (60 days, Wilcoxon, 10 seeds) reveals a cooperative survival network: 9 of 15 mechanisms become individually critical (ΔQΔQ up to −93.7%-93.7\%), while moderate conditions mask individual contributions. Sim-Selection sleep adds 37% stability with 47.4% storage reduction (p≤5.1×10−3p \le 5.1\times 10^{-3}); TripleCopyMemory retains S(t)=0.912S(t)=0.912 at 30 days; multi-layer routing beats a flat baseline by +20.7%+20.7\% F1 on LoCoMo, +19.5%+19.5\% on MemoryArena. A cross-provider bias-direction check (ΔGPT-Anth=−0.0001Δ_\text{GPT-Anth}=-0.0001 for ZB vs. −0.049-0.049 for Mem0) rules out LLM-judge-specific confounds. Open-source with 11,589 CI tests.
Apr 22, 2026cs.NE

Learning Hippo: Multi-attractor Dynamics and Stability Effects in a Biologically Detailed CA3 Extension of Hopfield Networks

We present a biologically detailed extension of the classical Hopfield/Marr auto-associative memory model for CA3, implementing ten populations (two asymmetric pyramidal subtypes, eight GABAergic interneuron classes), forty-seven compartments, multi-rule plasticity (recurrent Hebb, BCM anti-saturation, mossy-fiber short-term, endocannabinoid iLTD, burst-gated Hebb), and a bimodal cholinergic encoding/consolidation cycle. Evaluated on pattern completion across auto-associative, associative, and temporal regimes, and on a controlled inhibitory-proportion manipulation at N=256N{=}256, the full architecture exhibits \emph{three qualitative signatures absent from a minimal Hopfield baseline}: (i)~multi-attractor cross-seed behaviour at K=5K{=}5 with biologically realistic inhibitory proportions, where two of five seeds converge to positive attractors with margin +0.10−0.22{+}0.10{-}0.22 (Cohen's d=0.71d{=}0.71, one-sided p=0.08p{=}0.08); (ii)~target-selective associative recall in paired (A,B)(A, B) memory at K≥5K{\geq}5, where the full model retrieves BB from a partial cue of AA while the minimal model echoes AA (Pearson margin Δ=+0.163Δ{=}{+}0.163 at K=5K{=}5); (iii)~reduced cross-seed variance of the full model below the minimal baseline under clean upstream, with ratios 1.0−3.01.0{-}3.0. These three signatures are architecture-specific: they appear consistently across independent regimes and are absent from the minimal control.
Apr 21, 2026cs.LG

On the Conditioning Consistency Gap in Conditional Neural Processes

Neural processes are meta-learning models that map context sets to predictive distributions. While inspired by stochastic processes, NPs do not generally satisfy the Kolmogorov consistency conditions required to define a valid stochastic process. This inconsistency is widely acknowledged but poorly understood. Practitioners note that NPs work well despite the violation, without quantifying what this means. We address this gap by defining the conditioning consistency gap, a KL divergence measuring how much a conditional neural process's (CNP) predictions change when a point is added to the context versus conditioned upon. Our main results show that for CNPs with bounded encoders and Lipschitz decoders, the consistency gap is O(1/n2)O(1/n^2) in context size nn, and that this rate is tight. These bounds establish the precise sense in which CNPs approximate valid stochastic processes. The inconsistency is negligible for moderate context sizes but can be significant in the few-shot regime.
Apr 20, 2026cs.CL

On the Emergence of Syntax by Means of Local Interaction

Can syntactic processing emerge spontaneously from purely local interaction? We present a concrete instance on a minimal system: an 18,658-parameter two-dimensional neural cellular automaton (NCA), supervised by nothing more than a 1-bit boundary signal, is trained on the membership problem of an arithmetic-expression grammar. After training, its internal L×LL \times L grid spontaneously self-organizes into an ordered, spatially extended representation that we name Proto-CKY. This representation satisfies three operational criteria for syntactic processing: expressive power beyond the regular languages, structural generalization beyond the training distribution, and an internal organization quantitatively aligned with grammatical structure (Pearson r≈0.71r \approx 0.71). It emerges independently on four context-free grammars and regenerates spontaneously after perturbation. Proto-CKY is functionally aligned with the CKY algorithm but formally distinct from it: it is a physical prototype, a concrete instantiation of a mathematical ideal on a physical substrate, and the systematic distance between the two carries information about the substrate itself.
Mar 23, 2026cs.AI

AI Mental Models: Learned Intuition and Deliberation in a Bounded Neural Architecture

This paper asks whether a bounded neural architecture can exhibit a meaningful division of labor between intuition and deliberation on a classic 64-item syllogistic reasoning benchmark. More broadly, the benchmark is relevant to ongoing debates about world models and multi-stage reasoning in AI. It provides a controlled setting for testing whether a learned system can develop structured internal computation rather than only one-shot associative prediction. Experiment 1 evaluates a direct neural baseline for predicting full 9-way human response distributions under 5-fold cross-validation. Experiment 2 introduces a bounded dual-path architecture with separate intuition and deliberation pathways, motivated by computational mental-model theory (Khemlani & Johnson-Laird, 2022). Under cross-validation, bounded intuition reaches an aggregate correlation of r = 0.7272, whereas bounded deliberation reaches r = 0.8152, and the deliberation advantage is significant across folds (p = 0.0101). The largest held-out gains occur for NVC, Eca, and Oca, suggesting improved handling of rejection responses and c-a conclusions. A canonical 80:20 interpretability run and a five-seed stability sweep further indicate that the deliberation pathway develops sparse, differentiated internal structure, including an Oac-leaning state, a dominant workhorse state, and several weakly used or unused states whose exact indices vary across runs. These findings are consistent with reasoning-like internal organization under bounded conditions, while stopping short of any claim that the model reproduces full sequential processes of model construction, counterexample search, and conclusion revision.
Mar 16, 2026cs.LG

The Metric Slingshot: Navigational Reuse as Width-Optimal Structural Decoupling in Continual Learning

The mammalian brain, most extensively studied in rodents and bats, solves an enormous variety of non-spatial cognitive tasks using neural circuitry, including grid cells, place cells, and hippocampal indexing, that originally evolved for physical navigation. We formalize the above observation within the local Urysohn width (LUW) framework for continual learning. The central construct is the \emph{metric slingshot}: a learned embedding φ:X→Zφ: X \to Z that maps an arbitrary learning problem into a navigational latent space ZZ where pre-evolved contraction maps (grid cells) already provide the metric machinery, so that only the topological indexing subproblem must be solved de novo. We prove three results. First, the optimal spacing of multi-scale grid cell modules is a geometric series whose ratio is determined by the Ω(wlog⁡w)Ω(w \log w) sample complexity bound of the LUW framework; for ecologically plausible parameters, optimality yields r∗≈1.4r^* \approx 1.4--1.71.7, matching electrophysiological measurements in rodent medial entorhinal cortex. Second, the slingshot preserves the width hierarchy with a Lipschitz-controlled transfer bound: a contractive embedding into a fine-resolution navigational space reduces the effective number of contexts the learner must discover. Third, the anatomical separation of the ventral (what'') and dorsal (where'') visual streams achieves the structural decoupling required by Metric-Topology Factorization (MTF) \emph{by architecture}, without gradient-routing mechanisms. We demonstrate that such metric slingshot is applicable to both perception cognition and motor control. Together, these results provide a unified, complexity-theoretic account of navigation in non-spatial domain, grid cell multi-modularity, and hippocampal-neocortical complementary learning as consequences of a single exaptation principle: metric slingshot.
Feb 27, 2026cs.SD

SHINE: Sequential Hierarchical Integration Network for EEG and MEG

How natural speech is represented in the brain constitutes a major challenge for cognitive neuroscience. Reconstructing the speech envelope and Mel spectrogram from EEG and MEG provides a time-resolved way to study its temporal and spectral structure. Speech-related neural activity spans sensors and temporal scales; extracting these representations while adapting the use of context to each acoustic target is a central problem in speech reconstruction. We propose SHINE, a Sequential Hierarchical Integration Network for EEG and MEG. A residual sensor adapter unifies input dimensions, intermediate dilated-block states retain temporal depth, and a target- and time-dependent gate fuses local hierarchical and attention-enhanced context predictions. Across two EEG and two MEG datasets, SHINE has the highest mean envelope and mean-Mel Pearson correlations among nine local baseline implementations on all eight dataset-metric combinations. SHINE also placed second in the speech-detection Extended Track of the NeurIPS 2025 PNPL Competition. Code will be released at https://github.com/xuxiran/SHINE.
Nov 18, 2025q-bio.NC

Teaching signal synchronization in deep neural networks with prospective neurons

Working memory requires the brain to maintain information from the recent past to guide ongoing behavior. Neurons can contribute to this capacity by slowly integrating their inputs over time, creating persistent activity that outlasts the original stimulus. However, when these slowly integrating neurons are organized hierarchically, they introduce cumulative delays that create a fundamental challenge for learning: teaching signals that indicate whether behavior was correct or incorrect arrive out-of-sync with the neural activity they are meant to instruct. Here, we demonstrate that neurons enhanced with an adaptive current can compensate for these delays by responding to external stimuli prospectively -- effectively predicting future inputs to synchronize with them. First, we show that such prospective neurons enable teaching signal synchronization across a range of learning algorithms that propagate error signals through hierarchical networks. Second, we demonstrate that this successfully guides learning in slowly integrating neurons, enabling the formation and retrieval of memories over extended timescales. We support our findings with a mathematical analysis of the prospective coding mechanism and learning experiments on motor control tasks. Together, our results reveal how neural adaptation could solve a critical timing problem and enable efficient learning in dynamic environments.
Jul 14, 2025cond-mat.dis-nn

Dynamical stability for dense patterns in attractor neural networks

Recurrent neural networks are canonical models of biological memory. In these models, memories are represented by distributed patterns of neural activity that are stored in the recurrent connections between neurons, such that they become attractors of the network's dynamics. During memory recall, network dynamics thus converge toward one of these memory patterns when started from a noisy or partial cue. Therefore, memory performance critically hinges on the dynamical stability of the stored patterns. However, previous theoretical approaches only studied dynamical stability under highly restrictive conditions that do not readily apply to biological neural circuits. Here, we develop a theory of the local stability of discrete fixed points in a broad class of networks with graded neural activities and in the presence of noise. Using methods from random matrix theory, we analyze the bulk and outliers of the eigenvalue spectra of the Jacobians that characterize network dynamics around fixed points. We show that either all fixed points are stable or all of them are unstable, depending on whether their number is below a ``critical load for stability'', which is distinct from the classical critical capacity that measures the maximal number of achievable fixed points regardless of their stability. We further analyze the dependence of this critical load for stability on experimentally measurable quantities characterizing the statistics of memory patterns and the activation functions of neurons. Our analysis highlights the computational benefits of sparse-like patterns and threshold-linear activation functions and offers testable predictions for neural circuits supporting memory.
Feb 26, 2024q-bio.NC

ELiSe: Efficient Learning of Sequences in Structured Recurrent Networks

Behavior can be described as a temporal sequence of actions driven by neural activity. To learn complex sequential patterns in neural networks, memories of past activities need to persist on significantly longer timescales than the relaxation times of single-neuron activity. While recurrent networks can produce such long transients, training these networks is a challenge. Learning via error propagation confers models such as FORCE, RTRL or BPTT a significant functional advantage, but at the expense of biological plausibility. While reservoir computing circumvents this issue by learning only the readout weights, it does not scale well with problem complexity. We propose that two prominent structural features of cortical networks can alleviate these issues: the presence of a certain network scaffold at the onset of learning and the existence of dendritic compartments for enhancing neuronal information storage and computation. Our resulting model for Efficient Learning of Sequences (ELiSe) builds on these features to acquire and replay complex non-Markovian spatio-temporal patterns using only local, always-on and phase-free synaptic plasticity. We showcase the capabilities of ELiSe in a mock-up of birdsong learning, and demonstrate its flexibility with respect to parametrization, as well as its robustness to external disturbances.
Date pendingcs.NE

A Gradient-based yet Spike-Timing-Dependent Solution to the Feedback Learning Problem in Neural Microcircuits

The brain uses discrete spikes for dynamic computation, yet, how neural microcircuits (NMCs) solve temporal credit assignment using local spike timing remains a fundamental open question. Dominant spiking neural network (SNN) approaches circumvent this by approximating backpropagation through surrogate gradients, decoupling learning from biological spike timing. Here, we reformulate temporal credit assignment as a state separation problem: extracting task-required components induced by historical perturbations directly from the current neural state. This enables an online feedback learning framework for NMCs through a gradient tunneling (GT) algorithm and the lead-lag expansion technique that derives credit assignment from local synaptic spike timing, while remaining compatible with ANN-SNN hybrid architectures. Experimentally, GT-trained NMCs excel at long-timescale evidence integration and noise-robust memory retention, and perform comparably to leading SNN online learning methods on real-world benchmarks with far fewer parameters. The proposed framework addresses the two-decade-old NMC feedback learning problem and suggests a computationally plausible explanation for the brain's learning mechanisms.