Compute-in-Memory

Momentum

5 papers in the last four weeks, against 2 the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 34

Feb 24, 2026cs.LG

Dynamic Symmetric Point Tracking: Tackling Non-ideal Reference in Analog In-memory Training

Analog in-memory computing (AIMC) performs computation directly within resistive crossbar arrays, offering an energy-efficient platform to scale large vision and language models. However, non-ideal analog device properties make the training on AIMC devices challenging. In particular, its update asymmetry can induce a systematic drift of weight updates towards a device-specific symmetric point (SP), which typically does not align with the optimum of the training objective. To mitigate this bias, most existing works assume the SP is known and pre-calibrate it to zero before training by setting the reference point as the SP. Nevertheless, calibrating AIMC devices requires costly pulse updates, and residual calibration error can directly degrade training performance. In this work, we present the first theoretical characterization of the pulse complexity of SP calibration and the resulting estimation error. We further propose a dynamic SP estimation method that tracks the SP during model training, and establishes its convergence guarantees. In addition, we develop an enhanced variant based on chopping and filtering techniques from digital signal processing. Numerical experiments demonstrate both the efficiency and effectiveness of the proposed method.
Nov 19, 2025cs.AR

DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory Architectures

High-performance Host processors can integrate Processing-In-Memory (PIM) devices, which can accelerate memory-intensive kernels of Machine Learning (ML) models, including Large Language Models (LLMs), by leveraging the large memory bandwidth available at PIM cores. However, Host processor needs consecutive elements distributed across DRAM banks, while PIM cores need consecutive elements within their local banks. This necessitates data rearrangements in ML kernel execution that pose significant performance and programmability challenges, further exacerbated by the need to support diverse PIM devices. Current compilation approaches lack systematic optimization for diverse ML kernels and multiple PIM devices, and may largely ignore data rearrangement costs during the compute code optimization step. We show that data rearrangements and compute code optimization are interdependent, and need to be jointly optimized during the tuning process. Therefore, we design DCC, the first data-centric ML compiler for PIM systems that jointly co-optimizes data rearrangements and compute code in a unified tuning process. DCC integrates a multi-layer PIM abstraction to support multiple PIM backends. DCC enables effective co-optimization of data partitioning strategies with compute loop partitioning schemes. DCC applies PIM-specific code optimizations, and leverages a fast and accurate performance prediction model to select the bestperforming code schedule for a given kernel on a target PIM architecture. Our evaluations in various individual ML kernels show that DCC achieves up to 7.68x speedup (2.21x average) on HBM-PIM, and up to 13.17x speedup (3.92x average) on AttAcc PIM, over GPU-only execution. In end-to-end LLM inference, DCC on AttAcc accelerates GPT-3 and LLaMA-2 by 4.52x average (up to 7.71x in LLaMA-2) over GPU. DCC is open-sourced at https://github.com/SPIN-Research-Group/DCC.
Aug 16, 2025cs.LG

FedUHD: Unsupervised Federated Learning using In-Memory Hyperdimensional Computing

Unsupervised federated learning (UFL) enables privacy-preserving distributed training without data labeling, yet practical deployment remains challenging due to non-IID data, high computational and communication costs at edge devices, and sensitivity to communication noise. We propose FedUHD, the first UFL framework based on Hyperdimensional Computing (HDC). On the client side, FedUHD employs kNN-based cluster hypervector removal to mitigate non-IID effects by filtering detrimental local outliers. On the server side, cluster-aware HDC aggregation leverages cluster-level statistics to stabilize learning across heterogeneous clients. To further improve efficiency, we design a compute-in-memory (CIM) accelerator based on a novel phase-change memory (PCM) device, integrated with lightweight ASIC digital modules to execute the client-side HDC pipeline within the accelerator. The intrinsic robustness of HDC to low precision and device variations enables efficient mapping onto analog PCM crossbars, exploiting massive parallelism while minimizing data movement. Experimental results show that FedUHD achieves comparable accuracy to state-of-the-art neural network-based UFL methods across all datasets. On HAR and CIFAR10/100, FedUHD delivers an average 2,239x speedup and 1,542x higher energy efficiency on GPU. In addition, FedUHD reduces communication cost by up to 176x on HAR and CIFAR10/100 and demonstrates greater robustness than Orchestra under communication noise. Compared to GPU implementation of FedUHD, the proposed PCM-based accelerator provides an additional 4.07x speedup and three orders of magnitude higher energy efficiency on average. Furthermore, the results demonstrate the benefit of PCM over RRAM as a CIM substrate.
May 20, 2025cs.ET

Optimizing Binary and Ternary Neural Network Inference on RRAM Crossbars using CIM-Explorer

Using Resistive Random Access Memory (RRAM) crossbars in Computing-in-Memory (CIM) architectures offers a promising solution to overcome the von Neumann bottleneck. Due to non-idealities like cell variability, RRAM crossbars are often operated in binary mode, utilizing only two states: Low Resistive State (LRS) and High Resistive State (HRS). Binary Neural Networks (BNNs) and Ternary Neural Networks (TNNs) are well-suited for this hardware due to their efficient mapping. Existing software projects for RRAM-based CIM typically focus on only one aspect: compilation, simulation, or Design Space Exploration (DSE). Moreover, they often rely on classical 8 bit quantization. To address these limitations, we introduce CIM-Explorer, a modular toolkit for optimizing BNN and TNN inference on RRAM crossbars. CIM-Explorer includes an end-to-end compiler stack, multiple mapping options, and simulators, enabling a DSE flow for accuracy estimation across different crossbar parameters and mappings. CIM-Explorer can accompany the entire design process, from early accuracy estimation for specific crossbar parameters, to selecting an appropriate mapping, and compiling BNNs and TNNs for a finalized crossbar chip. In DSE case studies, we demonstrate the expected accuracy for various mappings and crossbar parameters. CIM-Explorer can be found on GitHub.