Post-training foundation video models on heterogeneous reward-weighted data usually assume that all data categories induce compatible updates. This assumption is fragile when categories correspond to different skills, domains, or evaluation dimensions. We study this problem in text-to-video post-training, where VBench2.0 dimensions define data buckets and an external multimodal reward pipeline assigns sample weights. We propose G3-LoRA (Gradient-Guided Grouped LoRA), a data organization procedure that probes category-level gradients induced by reward-weighted video samples, removes the shared global update direction, clusters categories by residual gradient compatibility, trains group-specific LoRA experts, and consolidates them into one adapter by weight merging followed by on-policy distillation from the experts. We motivate this procedure by viewing reward-weighted flow matching as velocity-field regression: incompatible reward dimensions may prefer different denoising directions in overlapping noisy latent regions, causing shared LoRA training to average capabilities. On Wan2.1-T2V-1.3B-Diffusers, the merged grouped adapter improves the matched VBench2.0 evaluation over the base model, a joint reward-weighted LoRA baseline, and random, semantic, and raw-gradient partitions trained with the same pipeline; an independent evaluator agrees, and on CogVideoX-2B grouping avoids the negative transfer of joint training. The gain is not uniform: merging compresses the largest specialist gains, distillation recovers part of this loss, and camera motion and several local-quality dimensions remain challenging. Together, these results suggest that gradient compatibility can serve as a practical diagnostic for organizing reward-weighted video post-training data.
Figures & tables
Figure 1: Method overview. G 3 -LoRA treats post-training categories as optimization objects. It uses reward-weighted category gradients to estimate compatibility, clusters categories after removing the global mixed direction, trains group-specific LoRA experts, merges them, and distills the experts back into the merged adapter.
Figure 2: Residual gradient compatibility across VBench2.0 dimensions. Each entry is the mean cosine similarity between two category gradients after projecting out the global mixed-gradient direction. Dimensions are ordered by the final M=5 grouping, so block structure indicates compatible categories and blue off-block regions indicate potential interference.
Method
Mean
Δ Mean
Improved / degraded dims
Worst drop vs. Base
Base Wan2.1
47.07
–
–
–
Joint LoRA
46.99
-0.07
9 / 8
-8.00
Random partition A
46.38
-0.69
6 / 9
-8.21
Random partition B
45.82
-1.25
8 / 8
-16.00
Semantic partition
47.20
+0.13
–
–
Raw-gradient partition
47.35
+0.28
–
–
Table 1: VBench2.0 comparison on Wan2.1. Grouped LoRA (Ours) obtains the best mean among merged adapters and a smaller worst drop than joint LoRA or the random partitions; OPD raises the mean further. Ties count as neither improved nor degraded; “–”: not computed.
Figure 3: Per-dimension score changes over the base model. Grouped LoRA (Ours) improves multi-view consistency and motion order, while material, human interaction, and camera motion reveal remaining dimension-specific tradeoffs. † Group-only experts are a non-deployable diagnostic: each dimension is scored with the expert of its own group. Colors use a fixed range of −20 to +20 ; the upper colorbar extension indicates values above +20 . Annotations show score changes rounded to one decimal place; the last column is the mean change across 17 dimensions.
Figure 4: Qualitative comparison on representative VBench2.0 dimensions. We compare Base Wan2.1 and Grouped LoRA (Ours) on Composition and Thermotics prompts; each row shows five temporal frames from one generated video. In the Composition case, Grouped LoRA more consistently preserves the relation among the dog, cat, and basket. In the Thermotics case, it shows clearer signs of cheese softening or melting. These examples are qualitative case studies and complement, rather than replace, the automatic VBench2.0 scores.
CogVideoX-2B (VBench2.0)
Wan2.1 (VideoScore-v1.1)
Base
39.87
Base
2.914
Joint LoRA
27.95
Grouped LoRA (Ours)
3.080
Grouped LoRA (Ours)
40.99
Ours − Joint
+13.04 [9.08, 16.92]
Ours − Base
+0.166 [0.140, 0.192]
Table 2: Second backbone and independent evaluator. Left: CogVideoX-2B, VBench2.0 macro score over 17 dimensions ( ×100 ). Right: VideoScore-v1.1 five-aspect mean on the same Wan2.1 videos used in Table 1 . Brackets: paired 95% bootstrap CI.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Value
Base model
Wan2.1-T2V-1.3B-Diffusers
Scheduler
UniPC
Steps
50
Guidance scale
5.0
Sample/flow shift
5.0
Resolution
832×480
Appendix
Table 3: Generation parameters for the quick VBench2.0 comparisons.
Item
Setting
Trainable modules
q, k, v, output projection LoRA
Objective
reward-weighted flow matching MSE
Reward normalization
sample weight / global mean weight
Text length
512 tokens
Probe samples
100 per dimension for the full gradient-probe run
Probe stages
0/25/50/75/100%
Appendix
Table 4: LoRA post-training parameters.
Item
Setting
Student initialization
merged rank-40 adapter ϕ0 (student rank 40)
Anchor
frozen merged adapter ϕ0
Teachers
five frozen group experts, routed by dimension
Prompts
training prompts of the reward pool (videos, weights unused)
Loss weights
λOPD=1.0 , λ=0.1 (Eq. 6 )
Student rollout
8 sampler steps from fresh noise, no gradient
Appendix
Table 5: On-policy distillation (OPD) settings on Wan2.1.
Stage
Hardware
GPU-hours
Note
Joint training and gradient probes
NVIDIA A800
13.4
Five group experts
NVIDIA A800
33.5
OPD consolidation
NVIDIA RTX 4090
≈ 205
16 GPUs, 12 h 48 min
Appendix
Table 6: Compute. GPU-hours are measured on the listed hardware.
Item
Value
Produced reward records
9,512
Records retained for training
9,448 (99.33%)
Reward-pipeline failures
0
Duplicate video paths
0
QA-structure anomalies
0
Samples at a weight-clipping bound
0
Appendix
Table 7: Audit of the Wan2.1 reward pipeline.
Dimension
Base
Joint LoRA
Random A
Random B
Group-only
Grouped r8
Grouped r40 (Ours)
+OPD (Ours)
Human Anatomy
88.21
86.27
87.38
86.66
83.88
85.00
85.99
85.94
Human Clothes
97.92
97.96
95.92
97.96
92.00
83.33
95.92
98.40
Human Identity
73.37
71.69
73.40
73.01
68.54
77.15
71.59
71.49
Composition
48.10
46.30
49.40
50.20
53.60
46.30
47.80
47.30
Mechanics
56.52
61.70
54.35
61.70
72.92
70.00
57.78
58.49
Material
36.67
29.03
34.48
31.03
30.00
45.83
32.14
32.70
Appendix
Table 8: Dimension-level scores under the quick VBench2.0 evaluation protocol. Scores are percentages; ties are counted as neither improved nor degraded in aggregate tables. Grouped r8 is included as an appendix-only merge-rank diagnostic; the main comparison uses Grouped r40 (Ours). Group-only scores each dimension with the expert of its own group (non-deployable diagnostic). +OPD is Grouped r40 after on-policy distillation.
Figure 5: Hierarchical clustering over mean residual-gradient compatibility. The M=5 cut separates five groups with positive within-group residual similarity and avoids negative within-group pairs.
Figure 6: Stage-wise residual-gradient compatibility. Each heatmap is computed after projecting out the global mixed-gradient direction at the corresponding probe stage. The repeated block structure provides evidence that the grouping is not determined by a single noisy probe.
Signal
Negative pairs
Separation
VBench2.0 mean
Raw gradient
0/136
0.046
47.35
Direct subtraction
76/136
0.635
48.18 (same partition)
Projection residual (Eq. 4 )
75/136
0.633
48.18
Appendix
Table 9: Grouping signals on Wan2.1. Separation is the within-group minus between-group mean similarity of the resulting M=5 partition, computed on the corresponding matrix.
Groups
Within-group mean
Separation
Negative within-group pairs
M=4
0.345
0.559
3
M=5
0.473
0.633
0
Appendix
Table 10: Clustering diagnostics for the number of groups on Wan2.1.
Backbone / probe cohort
Negative pairs
Within-group
Between-group
Separation
Wan2.1 probes
75/136 (55.1%)
+0.473
−0.160
+0.633
CogVideoX-2B main probes
84/136 (61.8%)
+0.154
−0.102
+0.257
CogVideoX-2B cohort A
81/136 (59.6%)
+0.177
−0.121
+0.298
CogVideoX-2B cohort B
87/136 (64.0%)
+0.191
−0.121
+0.312
Appendix
Table 11: Residual-gradient conflict across backbones and probe cohorts.
Comparison
Pearson
Spearman
Sign agr. (all)
Sign agr. (strong)
Resamples vs. full reference
0.847
0.834
0.816
0.853
Across the three resamples
0.728
0.703
0.755
0.787
Appendix
Table 12: Stability of the projection-residual matrix under probe resampling on Wan2.1.
Relative step size
Spearman ρ
Permutation p
0.0005
0.256
0.0028
0.001
0.255
0.0038
0.002
0.252
0.0028
Appendix
Table 13: Residual similarity predicts held-out cross-dimension transfer on Wan2.1.
Scope
Base
Merged r40 (Ours)
Group-only
Merged – Group
+OPD (Ours)
All 17 dims
47.07
48.18
48.06
+0.12
49.91
Group 1
60.03
62.98
57.11
+5.87
64.11
Group 2
66.39
64.89
63.10
+1.79
66.41
Group 3
26.70
27.40
27.47
-0.07
27.37
Group 4
36.86
38.80
42.72
-3.92
42.10
Group 5
33.46
35.09
41.46
-6.37
37.98
Appendix
Table 14: Group-level retention summary. Group-only LoRAs are an oracle-style diagnostic evaluated only on assigned dimensions; Merged r40 (Ours) preserves the average while still compressing some group-specific skills. Bold indicates the highest score in each row.
Group
Dimension
Base
Merged r40 (Ours)
Group-only
Merged – Group
+OPD (Ours)
Group 1
Dynamic Spatial Relationship
34.00
38.00
32.00
+6.00
38.00
Group 1
Human Clothes
97.92
95.92
92.00
+3.92
98.40
Group 1
Instance Preservation
84.21
86.00
82.00
+4.00
86.00
Group 1
Motion Order Understanding
24.00
32.00
22.45
+9.55
34.04
Group 2
Human Anatomy
88.21
85.99
83.88
+2.11
85.94
Group 2
Human Identity
73.37
71.59
68.54
+3.05
71.49
Appendix
Table 15: Selected dimension-level retention results from the group-only LoRA evaluation. Scores are percentages under the quick VBench2.0 protocol.
Table 19: VideoScore-v1.1 on the matched Wan2.1 evaluation videos. The derived mean averages the five aspects; its paired 95% CI is [0.140,0.192] with P(Δ>0)=1.000 .