Diffusion Transformers with Mixture-of-Experts (MoE) routing are a leading recipe for scaling generative models. Classifier-Free Guidance (CFG) is essential for generation quality, yet excessively high guidance scales trigger collapse. We identify a previously unreported failure mode in their combination: the two CFG branches route independently, so their realized activations occupy different subspaces. The unconditional write then leaves the conditional subspace, and CFG amplifies that residual linearly in the guidance scale. We propose SAGE, a training-time regularizer that aligns unconditional MoE activations to the conditional subspace without restricting routing diversity, at zero inference cost. Toy experiments show that SAGE dramatically suppresses extreme drift by 9.2x. When scaled to a 1B-parameter text-to-image model, SAGE significantly improves generation quality, delivering a 9.3% boost in peak DPG-Bench performance. Extensive experiments demonstrate that SAGE consistently outperforms the baseline.
Figures & tables
Figure 1: Subspace alignment. Top: Conditional and unconditional hidden states hc and hu are routed independently, so the realized activations occupy different subspaces. (a) Standard CFG injects a leak δ(s) . (b) SAGE uses thin QR of the stacked conditional outputs Fc to yield an orthonormal basis Qc of VSc . Projecting Fu onto this subspace and penalizing the residual with Lsage enforces vu∈VSc at training time.
Figure 2: Toy experiment sample grids for Dense, MoE, and MoE+SAGE across s∈{1,4,7,15} , with per-class center drift ( Δd ; lower is better). The baseline MoE exhibits increasing drift at large s ; SAGE keeps the drift substantially smaller.
Excess Drift Δd(↓)
Mode Purity (↑)
Method
s=4
s=7
s=15
s=1
s=2
s=4
s=7
s=15
Dense
0.33 (0.9)
0.50 (2.0)
1.06 (4.3)
.993
.968
.794
.586
.501
MoE
0.72 (0.9)
1.83 (3.9)
5.59 (12.9)
.989
.929
.394
.121
.125
MoE + SAGE
0.34 (0.6)
0.39 (1.0)
0.39 (1.4)
.997
.989
.917
.781
.666
Table 1: Quantitative results on the toy experiment. Excess drift Δd reports the mean per-class distance increase relative to s=1 (matching the bar charts in fig. 2 ); the per-class maximum is shown in parentheses for cross-reference with the bar-chart annotations. Mode purity (higher is better) measures separation across the eight modes. MoE drift explodes at high guidance while MoE + SAGE stays close to Dense.
CFG guidance scale s
Benchmark
Method
1
3
5
7
10
Δ (peak → 10)
Δbase
DPG-Bench
Baseline
54.77
61.76
56.33
52.38
48.56
−13.20
–
SAGE ( λ=1 )
52.30
67.49
65.63
63.36
59.55
−7.94
+5.73
KL ( λkl=1 )
40.95
50.65
46.82
44.38
41.08
−9.57
−11.11
GenEval
Baseline
30.3
57.1
55.5
53.2
47.1
−10.0
–
SAGE ( λ=1 )
26.5
59.8
59.3
57.1
52.5
−7.3
+2.7
Table 2: Overall scores on DPG-Bench and GenEval across CFG scales. SAGE consistently outperforms the Baseline at all s≥3 and degrades more gracefully under high guidance, while the KL routing constraint hurts performance across the board. Δ (peak → 10) measures the change from the peak score to the score at s=10 ; Δbase shows the difference from the Baseline’s peak score. Bold denotes the best score.
Figure 3: Performance improvements of SAGE across CFG scales. SAGE (red) maintains high performance across the full guidance range, while the Baseline (blue) peaks at s=3 and declines sharply. The KL routing constraint (gray) consistently underperforms both methods, confirming that naive distributional alignment of routing decisions is counterproductive.
Figure 4: Qualitative comparison across CFG scales. Each panel shows the same prompt rendered by Baseline (top) and SAGE (bottom) at s∈{1,3,7,10,15} from left to right. At s=10 and 15 , the Baseline exhibits over-saturation, white patches, and structural collapse, while SAGE preserves coherent composition and realistic textures.
s=3
s=5
Method
DPG
GenEval
DPG
GenEval
Baseline
61.76
57.1
56.33
55.5
SAGE (from scratch)
67.49
59.8
65.63
59.3
Post-hoc SAGE (5k)
61.32
56.3
56.23
54.8
Post-hoc SAGE (10k)
60.20
55.3
55.73
53.5
Table 3: Training timing and loss weight analysis. Left: comparison of from-scratch SAGE and post-hoc fine-tuning of a pre-trained baseline for 5k or 10k steps, with λ=1 for all SAGE variants. Right: effect of the SAGE loss weight λ ; the default λ=1 performs best among the tested values on both benchmarks at both guidance scales.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Method
s
Global
Entity
Attribute
Relation
Other
Overall
Baseline
1
70.06
68.04
73.16
84.98
43.00
54.77
3
73.10
73.87
81.00
85.71
57.40
61.76
5
65.43
69.66
77.44
82.28
54.60
56.33
7
61.02
66.10
75.00
79.96
52.00
52.38
10
59.65
62.59
71.66
77.90
50.60
48.56
SAGE ( λ=1 )
1
69.45
65.94
71.04
84.91
42.30
52.30
Appendix
Table 4: DPG-Bench category scores (%) across CFG scales. Category scores aggregate raw per-question answers over all four images per prompt. Overall uses dependency-aware scoring and matches table 2 of the main paper; it is not the category average. Higher is better.
s=3
s=5
Category
Subcategory
Baseline
SAGE
KL
Baseline
SAGE
KL
Global
–
73.10
76.44
67.63
65.43
73.40
64.44
Entity
Whole
73.89
79.03
63.72
69.43
78.02
60.59
Part
77.25
80.08
71.58
75.44
81.01
71.09
State
71.91
75.53
65.63
67.66
74.29
63.20
Attribute
Color
87.90
88.52
84.71
85.32
87.87
82.35
Appendix
Table 5: Fine-grained DPG-Bench scores (%) at s∈{3,5} . All 13 second-level categories use the same four-image, pre-dependency aggregation as table 4 . SAGE and KL use λ=1 and λkl=1 , respectively. Higher is better.
Method
s
Single Obj.
Two Obj.
Counting
Colors
Position
Color Attr.
Overall
Baseline
1
50.0
18.2
20.3
44.1
25.3
23.8
30.3
3
79.4
46.0
40.0
79.3
51.5
46.8
57.1
5
82.5
43.9
40.3
75.8
48.8
41.8
55.5
7
79.4
43.9
40.3
68.6
49.8
37.3
53.2
10
73.1
42.4
33.1
58.2
42.5
33.3
47.1
SAGE ( λ=1 )
1
42.5
17.2
18.1
36.7
20.3
24.0
26.5
Appendix
Table 6: GenEval per-task accuracy (%) across CFG scales. Per-category GenEval scores for all three methods at every guidance scale s ; the “Overall” column is the mean over the six tasks and matches the GenEval rows of table 2 of the main paper.
s=3
s=5
Interval
DPG
GenEval
DPG
GenEval
Every step
67.49
59.8
65.63
59.3
Every 2 steps
67.21
59.9
65.84
59.1
Every 5 steps
66.73
59.4
65.12
58.8
Appendix
Table 7: Ablation: SAGE activation frequency ( s∈{3,5} ). DPG-Bench and GenEval scores for SAGE activation at every step, every second step, and every fifth step.
s=3
s=5
Configuration
DPG
GenEval
DPG
GenEval
8E + 1 shared, no SAGE
61.76
57.1
56.33
55.5
8E + 1 shared, SAGE ( λ=1 )
67.49
59.8
65.63
59.3
8E, no shared, no SAGE
57.82
53.0
52.41
51.4
8E, no shared, SAGE ( λ=1 )
64.05
56.4
62.18
55.6
ProMoE
63.12
58.4
57.95
56.8
Appendix
Table 8: Ablation: shared expert and ProMoE ( s∈{3,5} ). Comparison of MoE configurations with and without SAGE against ProMoE at both guidance scales.
Figure 5: Conditional vs. unconditional router similarity across CFG scales. SAGE (red) tracks the Baseline (blue) almost exactly on all four measures, whereas the KL routing constraint (gray, dashed) forces the two branches to agree (higher Jaccard/agreement, near-zero Jensen–Shannon (JS) divergence). Higher values indicate greater similarity for Top- 2 Jaccard, Top- 1 agreement, and routing-probability cosine similarity; lower values indicate greater similarity for JS divergence. Error bars are standard errors over 500 prompts and are smaller than the markers.
Figure 6: Conditional minus unconditional expert load per layer and expert at s=3 , averaged over 500 prompts and shown on a shared color scale. Baseline and SAGE exhibit comparably large conditional/unconditional load differences (routing diversity preserved), whereas the KL constraint suppresses them almost entirely (routing forced to agree).
Method
s=1
s=3
s=5
s=7
s=10
Baseline
13.90
18.54
26.11
33.11
45.38
SAGE ( λ=1 )
13.18
14.80
20.28
25.54
37.12
Δ (Base − SAGE)
+0.72
+3.74
+5.83
+7.57
+8.26
Appendix
Table 9: FID across CFG scales. Lower is better. Δ is Baseline − SAGE, positive values indicate a SAGE improvement.
Figure 7: Training dynamics of the Baseline and SAGE. (a) Flow-matching MSE loss; the inset magnifies a 4,000 -step interval and shows that the two runs remain closely matched. (b) MoE load-balancing auxiliary loss. (c) Percentage of dead experts. (d–f) Mean, per-layer maximum, and per-layer minimum of the subspace residual, respectively. SAGE preserves the primary optimization and routing-health statistics while consistently reducing the alignment residual.
College of Computing and Data Science, Nanyang Technological University, Singapore, 639798 · Institute of Artificial Intelligence of China Telecom (TeleAI), Shanghai, China, 200232