Rubric-Based Optimization for Text-to-Music Generation
Authors: Ping Wang, Guang Yang, Shao-Rong Su, Junkai Wu, Pang Wei Koh, Noah A. Smith
Organizations: Paul G. Allen School of Computer Science & Engineering, University of Washington · Department of Electrical and Computer Engineering, University of Washington · Allen Institute for AI
Post-training text-to-music generation requires reward signals that capture multiple aspects of musical quality beyond what any single automatic metric can measure. We study structured, rubric-based rewards from pretrained audio-language models (ALMs) as training signals for both autoregressive and diffusion-based music generators. An ALM scores each generated clip against the rubric; we rank candidates generated for the same text prompt by their scores and convert these rankings into preference pairs for DPO on both MusicGen-small and ACE-Step v1, and additionally use the rubric scores directly as scalar rewards for DiffusionNFT on ACE-Step v1. On MusicCaps, rubric-based optimization improves CLAP, SongEval, and Audiobox-Aesthetics simultaneously, with the strongest gains obtained by DiffusionNFT on ACE-Step. By contrast, on MusicGen-small, building preferences from any one of these automatic evaluators produces clear cross-metric trade-offs: the targeted evaluator improves while other independent evaluators deteriorate. We further study tempo, key, and instrumentation, where precise objective rewards are available. Directly optimizing these specialized rewards reliably improves the target attributes, whereas ALM rubrics provide only partial transfer for tempo and instrumentation and no measurable improvement for key. Together, these results suggest a practical division of labor: ALM rubrics are effective for broad perceptual qualities that are difficult to formalize, while specialized objective rewards remain preferable when reliable measurements are available.
Figures & tables
Figure 1: DPO on MusicCaps: change in each evaluator as a percentage of the untrained generator’s score (positive is better), for MusicGen-small (top row) and ACE-Step v1 (bottom row). Each panel is one DPO run, named by its preference source: the candidates were ranked either by a single automatic metric (hatched, left three panels) or by the ALM rubric of one rater (solid, right three panels). Within a panel the three bars are the three evaluators, so a run with no trade-off shows three positive bars. Each row has its own scale, and the largest bars in the top row pass through an axis break so that the small bars stay readable. Absolute deltas are in Appendix A.1 , Tables 1 and 2 ; untrained scores in Table 4 .
Figure 2: ACE-Step v1 under different post-training methods: DPO versus DiffusionNFT for each ALM rater. One panel per rater; within each panel, the three groups are DPO on rubric preferences (hatched; the ACE-Step rubric runs of § 5.1 ), DiffusionNFT with joint prompting (all rubric dimensions rated in one request), and DiffusionNFT with dimension-wise prompting (one request per dimension). The bars in each group are the three evaluators, with the same colours as Figure 1 , each shown as the percentage change of its score after post-training relative to the untrained model. All panels share one scale, with axis breaks for the largest Qwen2-Audio bars. Absolute deltas are in Appendix A.1 , Tables 2 and 3 ; untrained scores in Table 4 .
Figure 3: Objective measure of each attribute on the evaluation prompts as DiffusionNFT training proceeds on ACE-Step v1, for the rubric-reward condition (orange; Qwen3-Omni, joint prompting) and the objective-reward condition (blue), which optimizes the plotted measure directly; the dashed grey line is the untrained model. (a) Tempo: the mean tempo reward of Equation 9 , which is 1 when every clip is within ±5 BPM of the requested tempo. (b) Key: exact key accuracy (Equation 11 ); chance for 24 keys is 4.2% . (c) Instrumentation: the contrastive instrument score of Equation 12 , the requested instruments’ presence relative to the strongest non-requested one. Stars mark the peak of each curve within the plotted budget, labelled with its value; the plotted values are listed in Table 5 .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Preference metric
CLAP Δ
SongEval Δ
Audiobox Δ
MusicGen-small
CLAP
+0.0129
+0.0510
−0.1871
MusicGen-small
SongEval
+0.0054
+0.1951
−0.3453
MusicGen-small
Audiobox
−0.0186
−0.0529
+0.5842
ACE-Step v1
CLAP
−0.0043
+0.0049
−0.0077
ACE-Step v1
SongEval
−0.0021
−0.0019
−0.0052
ACE-Step v1
Audiobox
+0.0011
−0.0028
+0.0015
Appendix
Table 1: Change in CLAP, SongEval, and Audiobox-Aesthetics on the MusicCaps test captions after DPO with preferences ranked by a single automatic metric (rows), relative to the untrained generator; the left three panels of Figure 1 .
Model
Rating model
CLAP Δ
SongEval Δ
Audiobox Δ
MusicGen-small
Music Flamingo
+0.0055
+0.0104
+0.0354
MusicGen-small
Qwen2-Audio
+0.0051
+0.0222
+0.0315
MusicGen-small
Qwen3-Omni
+0.0013
+0.0031
+0.0311
ACE-Step v1
Music Flamingo
+0.0008
+0.0012
−0.0101
ACE-Step v1
Qwen2-Audio
+0.0033
+0.0129
−0.0054
ACE-Step v1
Qwen3-Omni
+0.0026
+0.0093
+0.0141
Appendix
Table 2: Change in CLAP, SongEval, and Audiobox-Aesthetics on the MusicCaps test captions after DPO with preferences ranked by the ALM rubric of each rater (rows), relative to the untrained generator; the right three panels of Figure 1 and the DPO groups of Figure 2 .
Rater
Prompting
CLAP Δ
SongEval Δ
Audiobox Δ
Music Flamingo
joint
−0.0392
−0.2249
−0.9210
Music Flamingo
dimension-wise
+0.0127
+0.0528
+0.1917
Qwen2-Audio
joint
−0.0576
+1.1983
+0.8848
Qwen2-Audio
dimension-wise
−0.0753
+0.6062
+0.5142
Qwen3-Omni
joint
+0.0080
+0.0204
+0.4183
Qwen3-Omni
dimension-wise
+0.0180
+0.1057
+0.3385
Appendix
Table 3: Change in CLAP, SongEval, and Audiobox-Aesthetics on the MusicCaps test captions after DiffusionNFT on ACE-Step v1 for each rater and prompting strategy, at the checkpoint after 125 updates, relative to the untrained generator; the DiffusionNFT groups of Figure 2 .
Untrained generator
CLAP
SongEval
Audiobox
MusicGen-small, 8 s clips
0.23
2.22 – 2.24
6.48 – 6.51
ACE-Step v1, 30 s clips, DPO runs
0.17 – 0.18
2.21
6.25
ACE-Step v1, 30 s clips, DiffusionNFT runs
0.22
1.96
6.32
Appendix
Table 4: Absolute scores of the untrained generators on the MusicCaps test captions, against which the deltas of Tables 1 – 3 are computed and to which the relative changes quoted in § 5.2 refer. Base clips are regenerated for each evaluation run with that run’s seeds, so the base score varies slightly across runs; ranges give the spread over the runs in the corresponding table.
Task (measure)
Condition
0
10
20
30
40
50
60
Tempo (tempo reward)
objective
0.268
0.435
0.523
0.545
0.555
0.524
0.500
rubric
0.277
0.289
0.245
0.270
0.290
0.343
0.353
Key (exact key accuracy)
objective
0.043
0.058
0.099
0.123
0.126
0.112
–
rubric
0.045
0.047
0.048
0.040
0.045
0.045
–
Instrument. (contrastive)
objective
0.414
0.498
0.592
0.538
–
–
–
rubric
0.414
0.436
0.443
0.402
–
–
–
Appendix
Table 5: Values plotted in Figure 3 : the objective measure of each attribute on the evaluation prompts after the given number of DiffusionNFT updates, for the rubric-reward and objective-reward conditions (bold: peak within the plotted budget). Update 0 is the untrained model. On instrumentation, the rubric condition also raises the direct requested-instrument probability (mean classifier probability of the requested instruments) from 0.246 to 0.287 at its peak.