cs.CLMar 28, 2026

Limited Stereotype Control Through Routing Reweighting in MoE Language Models

Authors: Junhyeok Lee, Han Jang, Kyu Sung Choi

Organizations: Interdisciplinary Program in Cancer Biology, Seoul National University College of Medicine · Department of Radiology, Seoul National University Hospital · Department of Radiology, Seoul National University College of Medicine · Healthcare AI Research Institute, Seoul National University Hospital

Abstract

Demographic prompts are routed differently from neutral prompts in Mixture-of-Experts (MoE) language models, motivating tests of routing-level stereotype control. We introduce Fairness-Aware Routing Equilibrium (FARE), a diagnostic framework combining demographic routing profiles, empirical layer selection, and fixed inference-time reweighting, and evaluate five MoE architectures in English. At the selected operating points, CrowS-Pairs preference changes by at most 1.3 percentage points; DeepSeekMoE selects no intervention. Paired 95% confidence intervals exclude decreases larger than 2.2 points on each intervened model, and the only nominally significant change (Qwen1.5, p = 0.015) does not survive multiple-comparison correction. OLMoE and Qwen3 nevertheless change nearly every top-k expert set. Noise controls, random and truncated synthetic profiles, and hard masking also move preference by at most 1.5 points. Four generation protocols on OLMoE, Mixtral, Qwen1.5, and Qwen3 show no consistent change in the measured toxicity, lexical, or reference-overlap metrics. Selection and evaluation items overlap, so these comparisons are not independent evaluations. The tested reweighting procedure offers limited stereotype control.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Safety-Oriented Routing Analysis of Mixtral MoE Under Benign and Harmful Prompts

    May 22, 2026Md Nurul Absar SiddikyMixture-of-Experts Language ModelsExpert Routing

  2. When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models

    May 8, 2026Youngsik Yoon, Siwei Wang, Wei Chen +1Mixture-of-Experts Language ModelsLLM Routing

  3. RA-MoE: Routing-Aligned Fine-Tuning for Multilingual Adaptation of Mixture-of-Experts Models

    May 27, 2026Guanzhi Deng, Kuan Wu, Haibo Wang +6Mixture-of-Experts Language ModelsLLM Fine-Tuning