cs.LGSep 28, 2026

Single-Layer MeMo as a Randomized Hamming-Kernel Classifier

Authors: Alessandro Straziota

Organizations: Department of Enterprise Engineering University of Rome “Tor Vergata” Rome, Italy

Abstract

MeMo (Zanzotto et al., 2025) is a recent language-model architecture that stores associations between token contexts and next tokens in a correlation matrix memory. In this work, we study its single-layer form and show that its ideal retrieval rule is a multiclass classifier based on the positional Hamming kernel. The MeMo architecture represents both the sequence features and the output labels with Gaussian random codes. Its score is therefore a doubly randomized sketch of the ideal classifier. Under independent input and output codebooks, we bound the errors introduced by context sketching and output decoding, characterize their dependence on model and data parameters, and give a margin-based guarantee for recovering the ideal prediction. Controlled simulations support the trends predicted by the analysis. On a restricted WikiText-2 next-token task, we compare single-layer MeMo with classical baselines and show that it can offer a useful trade-off among predictive accuracy, memory, and throughput, particularly on a GPU, where its matrix operations can be parallelized.

Figures & tables

Explore similar work

CardsList
  1. MeMo: Memory as a Model

    May 14, 2026Ryan Wei Heng Quek, Sanghyuk Lee, Alfred Wei Lun Leong +6Large Language Model Memory

  2. MoRE: Scaling mixture of experts with hardware-aware low-rank routing

    Sep 28, 2026Honam Wong, Surbhi Goel, Enric Boix-AdseràExpertsLarge Language Model Routing

  3. MoRE: Mixture of Reused Experts

    Sep 16, 2026Eric S. Qiu, Utku Umur Acikalin, Justin Lovelace +4Mixture-Of-ExpertsExperts