CORTEX: Learning to Share and Specialize in Dense Language Models
Organizations: Simon Fraser University · Southern University of Science and Technology · The University of British Columbia
Abstract
Large language models are trained on heterogeneous data mixtures, where different knowledge domains require both shared knowledge and specialization. Existing modular approaches typically impose explicit components or discover modules through interpretability analysis after training. In this work, we propose CORTEX, a learning dynamics-inspired framework that learns internal modularization within dense language models. CORTEX partitions trainable matrices into parameter groups and learns module assignments from domain-conditioned gradient and cross-domain gradient similarity. We introduce the selective lesion score and module-domain mutual information to characterize the target-domain lesion effects and alignment, and analyze how module assignment affects the trade-off between assignment bias and update magnitude. Experiments with 160M, Qwen3-8B, and Qwen3-32B backbone models show that CORTEX achieves the highest synthetic-domain exact match and largest average perplexity reduction, while remaining competitive on real-domain evaluations and forming identifiable modules.
Figures & tables
| Backbone / Metric | Approach | Domain 1 | Domain 2 | Domain 3 | Domain 4 | Domain 5 | Domain 6 |
| COPY | REVERSE | ADD | LOOKUP | LOOKUP ADD | REVERSE LOOKUP | ||
| 160M EM (%) | Best Baseline | (MoM) | (MoM) | (MoM) | (UpIT) | (MoM) | (Self-MoE) |
| CORTEX | |||||||
| General | Code | Math | Biomedical | Legal | Reasoning | ||
| Qwen3-8B NLL | Best Baseline | (UpIT) | (MoM) | (Self-MoE) | (UpIT) | (UpIT) | (Self-MoE) |
| CORTEX |
| Backbone | Approach | General | Code | Math | Biomedical | Legal | Reasoning |
| Qwen3-8B Downstream | Best Baseline | (UpIT) | (MoM) | (Self-MoE) | (UpIT) | (UpIT) | (Self-MoE) |
| CORTEX | |||||||
| Qwen3-32B Downstream | Best Baseline | (UpIT) | (MoE) | (MoM) | (MoE) | (UpIT) | (Self-MoE) |
| CORTEX |
| Variant | PPL | NLL | EM | |||||||
| COPY | REVERSE | ADD | LOOKUP | LOOKUP ADD | REV LOOKUP | |||||
| CORTEX w/o Shared | ||||||||||
| CORTEX | ||||||||||
| Backbone | Method | NLL | |||
| Target | Other | ||||
| Random | 0.4082 | 1.3860 | 0.0097 | 0.0086 | |
| 160M | CORTEX | 0.0734 | 1.4560 | 0.1901 | 0.0064 |
| Random | 1.9179 | 1.0238 | 0.0374 | 0.0108 | |
| Qwen3-8B | CORTEX | 1.7303 | 1.1534 | 0.1557 | 0.0106 |
| Random | 1.6571 | 1.1996 | 0.0298 | 0.0097 | |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Notation | Meaning |
| Knowledge-domain label. | |
| Input and output dimensions of the -th trainable weight matrix. | |
| Index set and the number of parameter groups in . | |
| Domain-conditioned gradient of the -th parameter group in the -th trainable parameter matrix w.r.t. the -th knowledge domain. | |
| Set and the number of trainable weight matrices in the backbone model. | |
| Set and the number of knowledge domains. |
| Hyperparameter | Synthetic 160M |
| Domains | 6 |
| Backbone | 160M decoder-only |
| Optimizer | AdamW |
| Selected learning rate | |
| Learning-rate candidates | |
| Mini-batch size / samples per update | 32 / 32 |
| Hyperparameter | Qwen3-8B |
| Domains | 6 |
| Backbone | Qwen3-8B |
| Optimizer | Adafactor |
| Selected learning rate | |
| Learning-rate candidates | |
| Mini-batch size / samples per update | 1 / 48 |
| Hyperparameter | Qwen3-32B |
| Domains | 6 |
| Backbone | Qwen3-32B |
| Optimizer | Adafactor |
| Selected learning rate | |
| Learning-rate candidates | |
| Mini-batch size / samples per update | 1 / 48 |
| Baseline | COPY | REVERSE | ADD | LOOKUP | LOOKUP ADD | REVERSE LOOKUP |
| Dense | ||||||
| Random | ||||||
| MoE | ||||||
| MoM | ||||||
| UpIT | ||||||
| Self-MoE |
| Baseline | General | Code | Math | Biomedical | Legal | Reasoning |
| Dense | ||||||
| Random | ||||||
| MoE | ||||||
| MoM | ||||||
| UpIT | ||||||
| Self-MoE |
| Baseline | General | Code | Math | Biomedical | Legal | Reasoning |
| Dense | ||||||
| Random | ||||||
| MoE | ||||||
| MoM | ||||||
| UpIT | ||||||
| Self-MoE |
| Backbone | General | Code | Math | Biomedical | Legal | Reasoning |
| Qwen3-8B | ||||||
| Qwen3-32B |
| Backbone | Original | Matched shuffle | Difference |
| 160M | |||
| Qwen3-8B |