cs.LGApr 2, 2025

AYLA: Architecting a loss landscape in shallow neural networks to accelerate feature recovery

Authors: Behnam Gheshlaghi, Shahin Atakishiyev

Abstract

Feature learning in shallow neural networks exhibits rich yet fragile dynamics, including prolonged plateaus, abrupt phase transitions, and sensitivity to optimization hyperparameters. While recent theoretical work has characterized these behaviors through the geometry of loss landscapes, saddle escape mechanisms, and emergent scaling laws, practical methods for actively shaping these dynamics remain limited. In this paper, we introduce AYLA, a principled loss reparameterization framework that dynamically modulates gradient magnitudes during training without altering the location of stationary points or optimal solutions. AYLA applies a smooth, sigmoid-controlled power-law transformation to empirical loss, yielding a state-dependent effective learning rate that accelerates descent in flat or saddle-dominated regions while stabilizing late-stage optimization. Crucially, AYLA preserves all critical points of the original objective, acting solely as a monotone transformation that reshapes optimization trajectories rather than objectives. We evaluate AYLA in controlled teacher student settings using two-layer tanh networks trained on synthetic Gaussian data. Across stochastic gradient descent and multiple loss-exponent schedules, AYLA consistently improves feature recovery. This evidence is observed in terms of weight alignment, per-neuron cosine similarity, hidden-activation correlation, and spectral properties of learned representations, while AYLA maintains competitive or faster loss convergence. Spectral analyses further demonstrate that AYLA mitigates rank collapse and promotes richer internal representations, signaling a transition from lazy to active feature-learning regimes. AYLA offers a lightweight, theoretically grounded way to improve shallow-network optimization, especially in resource-limited or noise-sensitive settings.

Figures & tables

Explore similar work

CardsList
  1. Removing spurious minima for planar features by skip connections

    Oct 1, 2026Jakob Paul Zimmermann, Moritz Grillo, Andrei Balakin +1Flat MinimaSkip

  2. Fast Learning Rate Transfer in Shallow Linear Networks at Growing Training Horizons

    Sep 28, 2026Mana Sakai, Masaaki ImaizumiBatch

  3. Scaling Laws from Sequential Feature Recovery: A Solvable Hierarchical Model

    May 14, 2026Arie Wortsman-Zurich, Hugo Tabanelli, Yatin Dandi +2Feature LearningScaling Laws