cs.LGOct 2, 2026

Conditional Capacity and Routing in Mixture-of-Experts Particle Transformers

Authors: Kaushik Pendiyala, Haris Zia, Trevin Lee, Timothy Legge, Alejandro J. De Leon, Zihan Zhao, Aaron Wang, Abhijith Gandrakota, +3 more

Organizations: University of California Davis, Davis, CA, USA · University of California San Diego, San Diego, CA, USA · University of Illinois at Chicago, Chicago, IL, USA · Fermi National Accelerator Laboratory, Batavia, IL, USA

Abstract

Mixture-of-Experts (MoE) models can increase parameter capacity without proportionally increasing active computation, but it is unclear how this trade-off behaves in particle-physics transformers. We study dense and MoE Particle Transformers on 188-class JetClass-II, varying expert count, routing capacity, top-K, and auxiliary loss. We find that, when token dropping is avoided, top-1 MoE models improve over the dense baseline at nearly unchanged nominal forward compute, while further increasing the number of stored experts produces little additional accuracy gain. Activating multiple experts per token yields additional predictive improvements at higher computational cost. Routing analyses show that expert assignments become more strongly associated with particle identity and kinematics in some configurations, but this structure does not increase monotonically with classification performance. These results highlight the need to distinguish stored parameter capacity, active computation, routing capacity, and routing organization when evaluating sparse expert models for jet classification. Code and experiment configurations are available at https://github.com/kpendiyala/MPT.

Figures & tables

Explore similar work

CardsList
  1. A theoretical model for task routing in mixture-of-expert transformers

    Jun 12, 2026Vinoth Nandakumar, Yongli Xiang, Yunzhi Yao +2Mixture ModelsTheory

  2. Distributionally Robust Mixture-of-Experts Training

    Oct 5, 2026Xin Teng, Muxiao Li, Hongyi WenBalanced Learning

  3. ProbMoE: Differentiable Probabilistic Routing for Mixture-of-Experts

    Jun 1, 2026Heng Zhao, Zilei Shao, Guy Van den Broeck +1Large Language Model RoutingProbabilistic Inference