Long-horizon LLM agents are frequently implemented using sparse mixture-of-experts (MoE) models, yet the co-design of agentic behavior and MoE structures remains underexplored. In this work, we comprehensively study the connections between agentic post-training and MoE expert selection. In off-the-shelf MoE models, we observe expert selection exhibits a specialized structure that naturally aligns with agentic trajectories. Specifically, expert routing overlaps more between turns where the agent performs semantically similar operations (e.g., READ, UPDATE) than between turns with differing operations. However, standard RL algorithms ignore this specialization, allowing the MoE routing to go uncontrolled during training, which empirically limit task performance and inference efficiency. To address this, we introduce a hierarchical routing control framework for agentic tasks. We explicitly encourage turn-level expert selections to align with agentic operations while regularizing token-level expert selections to maintain local consistency. To resolve stability issues that arise during post-training with the proposed methods, we further introduce an entropy-gated control mechanism. Overall, our routing control framework achieves over 10-point improvements in success rate on all evaluated benchmarks. These results demonstrate that agentic trajectory structure provides an effective signal for optimizing MoE capacity during RL post-training.
Figures & tables
Figure 1 : Overview of MoE routing control for agentic trajectories. Top: In each turn of interaction with the environment, agents generate thinking and tool-use fields, then receive environment feedback . Our routing control framework aims to align expert selection with agentic operations (e.g., READ , CREATE , and UPDATE ) in each turn, assigning distinct expert groups to different operations and encouraging similar expert selections within each turn. Bottom: With Qwen3-30B-A3B on AppWorld, our routing control framework (when used with standard RL) unlocks the performance limits and raises the inference throughput, by improving expert specialization (larger with-cross gap of expert usage similarity in Appendix A.1 ) and consistency (higher Jaccard scores).
Figure 2 : Default routing statistics of Qwen3-30B-A3B on AppWorld. Left: Pairwise cosine similarity of token-level expert distributions; the difference between within group and cross-group cosine similarity values indicates a relatively higher expert overlap between tokens that share the same operation label. Right: Jaccard score between token-level expert sets is relatively higher within the same field and turn indicating more overlapping expert routing patterns.
Figure 3 : Overview of our hierarchical routing control framework. Left: Turn-level control aggregates router distributions within each field and maximizes the mutual information between the routing distribution and the turn’s operation labels. Right: Token-level control measures the top- k expert-set gap between adjacent tokens and reinforces the preceding expert set only when this gap is small. The entropy gate is omitted here for clarity.
Figure 4 : Effect of entropy gating on RL with routing control. Crosses mark the first evaluation at which both splits reach zero TGC. Routing control alone induces training collapse, and RL training remains stable for all 200 steps when an appropriate entropy gate (target range [0,0.2] ) is applied.
Method
AppWorld
AutomationBench
Qwen3-30B-A3B-2507
Qwen3.5-35B-A3B
test-normal
test-challenge
TGC (%)
SGC (%)
TGC (%)
SGC (%)
Base Model
35.1
14.3
18.7
7.2
22.04
PPO
63.1
37.5
35.5
16.5
54.56
+ Ours
60.7 (2.4↓)
42.9 (5.4↑)
38.6 (3.1↑)
20.9 (4.3↑)
57.39 (2.83↑)
Table 1 : Agentic task performance (in percentages). Bold and underline mark the highest and second-highest scores per column. Arrows show percentage-point changes from the corresponding baseline. Our routing control framework improve the agentic RL on both benchmarks, supporting diverse RL algorithms.
Figure 5 : Routing statistics during AppWorld GRPO training. (a) Token-level control improve routing similarity. The numbers are averaged over all token pairs within fields. (b) Turn-level control increases the mutual information between expert selection and operation labels. (c) Turn-level control makes expert selections more separable by the operation labels. The reported numbers are averaged over all of the evaluation trajectories collected from the specified training step intervals. The error bars are the standard deviation of all data points.
Figure 6 : AppWorld rollout efficiency comparison under GRPO with turn-level control, token-level control, or both. Both controls improve inference throughput. The token-level control contributes the most to inference efficiency at the early training stage. The reported numbers are averaged over all of the evaluation trajectories collected from the specified training step intervals. The error bars are the standard deviations of all data points.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Routing behavior changes during training. Measured by the expert switches and expert-operation classification, our routing control framework improves the routing consistency and expert specialization.
Figure 8: Default routing statistics of Qwen3.5-35B-A3B on Automation. Left: Pairwise cosine similarity of token-level expert distributions; the difference between within group and cross-group cosine similarity values indicates a relatively higher expert overlap between tokens that share the same operation label. Right: Jaccard score between token-level expert sets is relatively higher within the same field and turn indicating more overlapping expert routing patterns.
Method
test-normal
test-challenge
TGC (%)
TGC (%)
Base Model
34.52
19.66
GRPO
66.70
38.80
+ routing_reg
70.24
43.17
+ switch_reg ( Yang et al., 2026 )
72.02
46.28
+ outlier_reg
0.0 (crashed)
0.0 (crashed)
Appendix
Table 2: Comparison of routing control variants on AppWorld, measured by task goal completion (TGC). The highest and second-highest scores in each column are bolded and underlined, respectively.
Method
Expert-Load CV ↓
AppWorld TGC (%) ↑
test-normal
test-challenge
Base Model
1.240
38.69
19.66
GRPO
1.241
69.64
46.28
+load-bal ( Shazeer et al., 2017 )
1.155
74.40
45.56
+token-level control & load-bal
1.321
69.05
43.17
+token-level & turn-level control (ours)
1.295
72.62
51.80
Appendix
Table 3: Load balancing and task performance on AppWorld. Load-balancing loss effectively reduces the CV of expert load, but this does not translate to better empirical performance.
Figure 9: Rollout framework overview. Each trajectory is handled by a separate process with its sandbox for environment execution. SGLang servers generate tokens in response to HTTP requests from trajectory processes.
Setting
AppWorld
AutomationBench
Model
Qwen3-30B-A3B-Instruct-2507
Qwen3.5-35B-A3B
Inference engine
SGLang
SGLang
Inference GPUs
8
8
Engine replicas
2
2
TP / EP per engine
4 / 4
4 / 4
Temperature
0.7
0.7
Appendix
Table 4: Evaluation inference settings. Sampling parameters apply at each assistant turn. TP and EP denote tensor and expert parallelism. a AppWorld uses 160 concurrent trajectories, except for PPO, which uses 80. b AutomationBench uses 144, except for PPO, which uses 40. Concurrency limits are totals across both engines. The explicit context cutoff is checked before each generation turn; an unset cutoff does not remove model or server context limits.
Trajectory
Generation Log-Likelihood
Default Routing
Fixed Routing
Δ (%)
positive ( R=1 )
-0.2778
-1.5897
140.51
negative ( R=0 )
-0.2795
-1.5292
138.19
Appendix
Table 5: Enforcing fixed expert routing substantially degrades trajectory likelihood. Generation log-likelihoods are computed from a set of example trajectories. Higher mean log-probability is better. Δ=2⋅(ℓfixed−ℓdefault)/(ℓfixed+ℓdefault) is the relative decrease.
Figure 10: Training curves for GRPO with direct router-entropy regularization. Direct entropy control causes an explosion in total action entropy and training collapse.
Figure 11: Controlled routing versus inference throughput. We manually collapse the routers’ weight matrices to ensure consistent expert selection. The efficiency gain from consistent routing primarily comes from parallelism across a batch of trajectories decoded simultaneously.
Turn Label
# Classes
test-normal
test-challenge
Random
5
66.67
39.33
Application Name
12
66.07
40.53
Operation
5
70.83
42.69
Appendix
Table 6: AppWorld TGC (%) at training step 199 for Qwen3-30B-A3B-Instruct-2507 trained with GRPO and turn-level routing control ( λMI=2×10−6 ). All three runs share the same saved training settings apart from the labeling schemes. Labeling each turn by its operations yields the best performance.
Figure 12: Histogram of expert shifts. After RL training with routing_reg (Appendix B.1 ) and switch_reg (Appendix B.2 ), high expert-shift counts consistently become less frequent.
Figure 13: Comparison of training curves with R3. On agentic tasks, R3 can stabilize the entropy dynamics, preventing potential explosion. However, this stability does not translate to better performance due to the loss of RL exploration.
Figure 14: Comparison of training curves with entropy loss. We confirm that entropy loss has no effect on the routers’ behavior or inference efficiency.
Threshold
Eval Step
test-normal
test-challenge
Collapse Step
0/8
199
64.88
35.97
–
2/8
199
61.31
39.09
–
4/8
199
66.67
42.21
–
6/8
159
0.00
0.00
149
8/8
139
0.00
0.00
129
Appendix
Table 7: Token-level routing control gap thresholds, evaluated using AppWorld TGC (%) with GRPO. Allowing too few expert shifts reduces our framework to standard RL; allowing too many expert shifts significantly destabilizes RL training.
Setting
test-normal
test-challenge
Entropy
Tokens/GPU/s
Shifted Experts
TGC (%)
TGC (%)
λturn=0 (GRPO)
66.67
38.85
0.1842
251.13
4.8393
λturn=10−3
0.00
0.00
0.2235
276.37
4.6192
λturn=10−4
0.60
0.96
0.1736
223.55
4.7340
λturn=10−5
63.69
36.93
0.1801
277.63
4.8370
λturn=10−6
63.69
41.73
0.1926
290.86
4.8425
Appendix
Table 8: Hyperparameter search for turn-level routing control. We apply only turn-level control to GRPO to determine the appropriate magnitude.
Setting
test-normal
test-challenge
Entropy
Tokens/GPU/s
Shifted Experts
TGC (%)
TGC (%)
λtoken=0 (GRPO)
64.88
35.97
0.1628
305.34
4.8263
λtoken=10−4†
0.00
0.00
0.2649
394.55
4.5584
λtoken=10−5†
0.00
0.00
0.2639
364.24
4.6042
λtoken=10−6‡
0.00
0.00
0.2010
295.37
4.7318
λtoken=10−7
66.67
42.21
0.1448
306.03
4.8191
Appendix
Table 9: Hyperparameter search for token-level routing control. We apply only token-level control to GRPO to determine the appropriate magnitude. TGC is reported at step 199 when available and at the last evaluation otherwise: † denotes step 59 and ‡ denotes step 79.
Setting
test-normal
test-challenge
Entropy
Tokens/GPU/s
Shifted Experts
TGC (%)
TGC (%)
GRPO
66.67
38.85
0.1842
251.13
4.8393
+ Turn-Level Control
62.50
42.45
0.1943
313.52
4.8132
+ Token-Level Control
67.86
41.73
0.2053
277.74
4.6676
+ Both
75.00
49.20
0.2331
312.94
4.7416
Appendix
Table 10: Individual and combined effects of routing controls on AppWorld. Turn-level control primarily improves model capability on hard tasks, while token-level control primarily improves inference throughput.
Figure 15: AppWorld rollout efficiency comparison under GRPO with turn-level control, token-level control, or both. (a) Both controls reduce the rollout time for a single trajectory. (b) Both controls improve inference throughput. The token-level control contributes the most to inference efficiency at the early training stage. The reported numbers are averaged over all of the evaluation trajectories collected from the specified training step intervals. The error bars are the standard deviation of all data points.
Reinforcement learning (RL) has become central to post-training of large language models. Recent advances in RL for Mixture-of-Experts (MoE) models have primarily focused on improving optimization stability and training efficiency, while treating the expert selection as a fixed component. Since routing determines the sparse computation paths that induce output distributions, expert selection offers an additional source of rollout diversity. Through empirical analysis, we find that perturbing expert routing effectively alters model output and increases rollout diversity, which is similar to increasing the decoding temperature. However, direct perturbation can activate unsuitable experts and substantially degrade rollout quality. Motivated by these observations, we introduce Expert-Space Exploration Reinforcement Learning (ESRL), an architecture-aware framework that explicitly explores the expert-routing space of MoE models. ESRL preserves high-confidence experts as anchors, and restricts stochastic routing to a plausible candidate pool, thereby retaining reliable computation paths. The perturbation strength is further adapted according to router entropy to avoid over-perturbation. To mitigate the routing mismatch introduced by perturbation, ESRL records the expert paths used during rollout and replays them during policy optimization. Experiments demonstrate that ESRL achieves the best performance across MoE backbones with top-K, top-1, and shared-expert routing, as well as across mathematics, science, and code tasks without additional sampling or computational cost. Specifically, ESRL on Qwen3-30B-A3B achieves the best among all compared methods, improving average Pass@1 and Pass@8 over GRPO by 3.2 and 4.5 percentage points, respectively. Further analyses of expert utilization and training dynamics provide insights into how exploiting MoE-specific routing structure benefits RL training.
Mixture-of-Experts (MoE) models have become a leading approach for decoupling parameter count from computational cost in large language models, yet effectively scaling MoE performance remains a challenge. Prior work shows that fine-grained experts enlarge the space of expert combinations and improve flexibility, but they also impose substantial routing overhead, creating a new scalability bottleneck. In this paper, we explore a complementary axis for scaling -- how expert outputs are aggregated. We theoretically show that replacing the standard weighted-summation aggregation with structural aggregation expands the expert-combination space without altering the experts or router, and enables possible multi-step reasoning within a single MoE layer. To this end, we propose DAG-MoE, a sparse MoE framework that employs a lightweight module to automatically learn the optimal aggregation structure among the selected experts. Extensive experiments under standard language modeling settings show that DAG-MoE consistently improves performance in both pretraining and fine-tuning, surpassing traditional MoE baselines.
Jiarui Feng, Hanqing Zeng, Karish Grover +11
Meta MRS · Washington University in St. Louis · Carnegie Mellon University +1
We present Mach-Mind-4-Flash, a 35B-parameter Mixture-of-Experts (MoE) agentic model with 3B activated parameters. Through post-training optimization alone without scaling pre-training compute, the model achieves performance on par with or surpassing that of 100B-parameter-class models. By introducing scalable agentic interaction environments for large-scale reinforcement learning, the model attains significant performance gains on real-world application tasks. Our pipeline comprises three stages: (1) a unified RL/OPD training infrastructure with dynamic multi-teacher scheduling and operator-level acceleration, delivering 17% end-to-end training speedup; (2) multiple domain-specific RL experts trained in parallel across Reasoning, General, and Agent tracks, then fused into a single generalist via Multi-Teacher On-Policy Distillation (MOPD) -- a routed reverse-KL objective that eliminates the see-saw degradation of mixed-reward RL; (3) Hybrid Median-length Policy Optimization (HMPO), a single-stage token-efficiency method that compresses reasoning chains by 19--46% with ≤0.7 percentage-point accuracy loss. Mach-Mind-4-Flash scores 92.70 on AIME'26, 82.82 on IFBench, 80.74 on Behavioral-SafetyBench, 75.80 on BFCL-v4, 72.31 on BrowseComp-zh, and 84.20 on ClawBench -- leading or matching models with 10--30× its activated size at a fraction of the inference cost.