cs.LGSep 28, 2026

SOLO: Pretraining Billion-Parameter Language Models with Shared-Output Local Learning

Authors: Bojian Yin, Shurong Wang, Yuqi Pan, Guoqi Li

Organizations: Institute of Automation, Chinese Academy of Sciences

Abstract

Large language models are trained with backpropagation, whose global gradient coordinates all layers but forces each to hold its activations and wait for the gradient to pass back through every deeper layer. Conventional local learning removes this update locking by training each module to predict the target through its own readout, but has not scaled to billion-parameter pretraining. We identify these private readouts as a key weakness, since they leave each module without information from deeper modules. We propose Shared-Output LOcal learning (SOLO), which replaces them with a shared, read-only copy of the final module's readout, the only one trained on the output of the whole network. Taken from the previous step, the copy transmits information from the final module without passing gradients between modules or reintroducing update locking. SOLO approaches backpropagation on Transformers of 340M to 2B parameters pretrained on 15B tokens, staying within one point in average zero-shot accuracy with a perplexity gap that narrows with scale. Readout ablations attribute SOLO's improvement over private readouts to sharing. Without update locking, each of p pipeline stages holds activations for O(1) micro-batches instead of O(p). The freed memory permits larger micro-batches, which reach up to 1.44x the best measured throughput of pipeline backpropagation on the same partition. To our knowledge, SOLO is the first local learning method to show such memory and throughput gains in billion-parameter language-model pretraining. Local learning thus becomes a practical alternative to backpropagation for large-scale pretraining.

Figures & tables

Appendix figures & tables26 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Local Support Learning

    Oct 1, 2026Assaf Ben-Kish, Akarsh Kumar, James Glass +1Catastrophic ForgettingLarge Language Model Memory

  2. Rethinking Local Learning: A Cheaper and Faster Recipe for LLM Post-Training

    May 6, 2026Hengyu Shi, Tianyang Han, Peizhe Wang +3Post-TrainingLayer-Wise

  3. NoLoCo: No-all-reduce Low Communication Training Method for Large Models

    Jun 12, 2025Jari Kolehmainen, Nikolay Blagoev, Semih Kara +3Large Language Model Training