cs.AISep 28, 2026

CORTEX: Learning to Share and Specialize in Dense Language Models

Authors: Chuiyang Meng, Ming Tang, Vincent W. S. Wong

Organizations: Simon Fraser University · Southern University of Science and Technology · The University of British Columbia

Abstract

Large language models are trained on heterogeneous data mixtures, where different knowledge domains require both shared knowledge and specialization. Existing modular approaches typically impose explicit components or discover modules through interpretability analysis after training. In this work, we propose CORTEX, a learning dynamics-inspired framework that learns internal modularization within dense language models. CORTEX partitions trainable matrices into parameter groups and learns module assignments from domain-conditioned gradient and cross-domain gradient similarity. We introduce the selective lesion score and module-domain mutual information to characterize the target-domain lesion effects and alignment, and analyze how module assignment affects the trade-off between assignment bias and update magnitude. Experiments with 160M, Qwen3-8B, and Qwen3-32B backbone models show that CORTEX achieves the highest synthetic-domain exact match and largest average perplexity reduction, while remaining competitive on real-domain evaluations and forming identifiable modules.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts

    Apr 20, 2026Jacob Morrison, Sanjay Adhikesaven, Akshita Bhagia +3Mixture-Of-ExpertsPost-Training

  2. Decodable But Not Detachable: Training Data Granularity Determines Parametric Modularity in Large Language Models

    Aug 10, 2026Marcus Armstrong, Navid Ayoobi, Arjun MukherjeeNeuronsSelectivity

  3. EMO: Pretraining Mixture of Experts for Emergent Modularity

    May 7, 2026Ryan Wang, Akshita Bhagia, Sewon MinMixture-Of-ExpertsExperts