cs.LGOct 2, 2026

Consideration Circuits: Depth Separation and Universality Beyond a Single Softmax

Authors: Junjie Xiao, Huiwen Jia

Organizations: Department of Mathematics, Peking University · Department of Industrial Engineering and Operations Research, University of California, Berkeley

Abstract

Most feature-based choice models, classical and deep, score items and apply a single softmax. We introduce consideration circuits (CC), feature-based models of multi-stage choice defined by directed acyclic graphs of multinomial logit (MNL) units. Source units assign probabilities to menu items, and internal units combine predecessor distributions using MNL weights computed from their probability-weighted feature summaries. On a three-item compromise task with fixed non-collinear features, menu-independent random-utility models (RUM), including a single MNL unit, suffer an error bounded away from zero. For CC, in contrast, we establish a sharp depth--norm separation: increasing depth from 22 to 33 reduces the optimal maximum taste-vector norm for error εε from Θ(log⁡(1/ε)/ε)Θ(\log(1/ε)/ε) to Θ(log⁡(1/ε))Θ(\log(1/ε)). The depth-22 lower bound holds for arbitrary width and menu-independent routing biases, while a five-node depth-33 circuit with zero routing biases attains the logarithmic rate. More generally, we characterize two geometric conditions that are necessary and sufficient for approximating arbitrary deterministic choice tables on finite menu families. Under these conditions, depth 33 suffices, while depth 44 achieves optimal logarithmic norm scaling whenever the family contains a non-singleton menu. In experiments, standalone tree circuits with fewer than 600600 parameters attain the lowest mean test negative log-likelihood (NLL) among the evaluated models on four fixed-pool benchmarks and the Expedia temporal split. As output heads, CC generalize the linear MNL readout and lower mean test NLL for every tested encoder on Expedia and Trivago.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits

    May 8, 2026Michael Li, Nishant SubramaniMechanistic InterpretabilityCircuits

  2. Input-Blind Controls Produce Substantial Oracle Headroom for Layer Programs in Multiple-Choice Evaluation

    Oct 7, 2026Yibei Guo, Rui Liu