Rubric-Based Optimization for Text-to-Music Generation
Authors: Ping Wang, Guang Yang, Shao-Rong Su, Junkai Wu, Pang Wei Koh, Noah A. Smith
Organizations: Paul G. Allen School of Computer Science & Engineering, University of Washington · Department of Electrical and Computer Engineering, University of Washington · Allen Institute for AI
Post-training text-to-music generation requires reward signals that capture multiple aspects of musical quality beyond what any single automatic metric can measure. We study structured, rubric-based rewards from pretrained audio-language models (ALMs) as training signals for both autoregressive and diffusion-based music generators. An ALM scores each generated clip against the rubric; we rank candidates generated for the same text prompt by their scores and convert these rankings into preference pairs for DPO on both MusicGen-small and ACE-Step v1, and additionally use the rubric scores directly as scalar rewards for DiffusionNFT on ACE-Step v1. On MusicCaps, rubric-based optimization improves CLAP, SongEval, and Audiobox-Aesthetics simultaneously, with the strongest gains obtained by DiffusionNFT on ACE-Step. By contrast, on MusicGen-small, building preferences from any one of these automatic evaluators produces clear cross-metric trade-offs: the targeted evaluator improves while other independent evaluators deteriorate. We further study tempo, key, and instrumentation, where precise objective rewards are available. Directly optimizing these specialized rewards reliably improves the target attributes, whereas ALM rubrics provide only partial transfer for tempo and instrumentation and no measurable improvement for key. Together, these results suggest a practical division of labor: ALM rubrics are effective for broad perceptual qualities that are difficult to formalize, while specialized objective rewards remain preferable when reliable measurements are available.
Figures & tables
Figure 1: DPO on MusicCaps: change in each evaluator as a percentage of the untrained generator’s score (positive is better), for MusicGen-small (top row) and ACE-Step v1 (bottom row). Each panel is one DPO run, named by its preference source: the candidates were ranked either by a single automatic metric (hatched, left three panels) or by the ALM rubric of one rater (solid, right three panels). Within a panel the three bars are the three evaluators, so a run with no trade-off shows three positive bars. Each row has its own scale, and the largest bars in the top row pass through an axis break so that the small bars stay readable. Absolute deltas are in Appendix A.1 , Tables 1 and 2 ; untrained scores in Table 4 .
Figure 2: ACE-Step v1 under different post-training methods: DPO versus DiffusionNFT for each ALM rater. One panel per rater; within each panel, the three groups are DPO on rubric preferences (hatched; the ACE-Step rubric runs of § 5.1 ), DiffusionNFT with joint prompting (all rubric dimensions rated in one request), and DiffusionNFT with dimension-wise prompting (one request per dimension). The bars in each group are the three evaluators, with the same colours as Figure 1 , each shown as the percentage change of its score after post-training relative to the untrained model. All panels share one scale, with axis breaks for the largest Qwen2-Audio bars. Absolute deltas are in Appendix A.1 , Tables 2 and 3 ; untrained scores in Table 4 .
Figure 3: Objective measure of each attribute on the evaluation prompts as DiffusionNFT training proceeds on ACE-Step v1, for the rubric-reward condition (orange; Qwen3-Omni, joint prompting) and the objective-reward condition (blue), which optimizes the plotted measure directly; the dashed grey line is the untrained model. (a) Tempo: the mean tempo reward of Equation 9 , which is 1 when every clip is within ±5 BPM of the requested tempo. (b) Key: exact key accuracy (Equation 11 ); chance for 24 keys is 4.2% . (c) Instrumentation: the contrastive instrument score of Equation 12 , the requested instruments’ presence relative to the strongest non-requested one. Stars mark the peak of each curve within the plotted budget, labelled with its value; the plotted values are listed in Table 5 .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Preference metric
CLAP Δ
SongEval Δ
Audiobox Δ
MusicGen-small
CLAP
+0.0129
+0.0510
−0.1871
MusicGen-small
SongEval
+0.0054
+0.1951
−0.3453
MusicGen-small
Audiobox
−0.0186
−0.0529
+0.5842
ACE-Step v1
CLAP
−0.0043
+0.0049
−0.0077
ACE-Step v1
SongEval
−0.0021
−0.0019
−0.0052
ACE-Step v1
Audiobox
+0.0011
−0.0028
+0.0015
Appendix
Table 1: Change in CLAP, SongEval, and Audiobox-Aesthetics on the MusicCaps test captions after DPO with preferences ranked by a single automatic metric (rows), relative to the untrained generator; the left three panels of Figure 1 .
Model
Rating model
CLAP Δ
SongEval Δ
Audiobox Δ
MusicGen-small
Music Flamingo
+0.0055
+0.0104
+0.0354
MusicGen-small
Qwen2-Audio
+0.0051
+0.0222
+0.0315
MusicGen-small
Qwen3-Omni
+0.0013
+0.0031
+0.0311
ACE-Step v1
Music Flamingo
+0.0008
+0.0012
−0.0101
ACE-Step v1
Qwen2-Audio
+0.0033
+0.0129
−0.0054
ACE-Step v1
Qwen3-Omni
+0.0026
+0.0093
+0.0141
Appendix
Table 2: Change in CLAP, SongEval, and Audiobox-Aesthetics on the MusicCaps test captions after DPO with preferences ranked by the ALM rubric of each rater (rows), relative to the untrained generator; the right three panels of Figure 1 and the DPO groups of Figure 2 .
Rater
Prompting
CLAP Δ
SongEval Δ
Audiobox Δ
Music Flamingo
joint
−0.0392
−0.2249
−0.9210
Music Flamingo
dimension-wise
+0.0127
+0.0528
+0.1917
Qwen2-Audio
joint
−0.0576
+1.1983
+0.8848
Qwen2-Audio
dimension-wise
−0.0753
+0.6062
+0.5142
Qwen3-Omni
joint
+0.0080
+0.0204
+0.4183
Qwen3-Omni
dimension-wise
+0.0180
+0.1057
+0.3385
Appendix
Table 3: Change in CLAP, SongEval, and Audiobox-Aesthetics on the MusicCaps test captions after DiffusionNFT on ACE-Step v1 for each rater and prompting strategy, at the checkpoint after 125 updates, relative to the untrained generator; the DiffusionNFT groups of Figure 2 .
Untrained generator
CLAP
SongEval
Audiobox
MusicGen-small, 8 s clips
0.23
2.22 – 2.24
6.48 – 6.51
ACE-Step v1, 30 s clips, DPO runs
0.17 – 0.18
2.21
6.25
ACE-Step v1, 30 s clips, DiffusionNFT runs
0.22
1.96
6.32
Appendix
Table 4: Absolute scores of the untrained generators on the MusicCaps test captions, against which the deltas of Tables 1 – 3 are computed and to which the relative changes quoted in § 5.2 refer. Base clips are regenerated for each evaluation run with that run’s seeds, so the base score varies slightly across runs; ranges give the spread over the runs in the corresponding table.
Task (measure)
Condition
0
10
20
30
40
50
60
Tempo (tempo reward)
objective
0.268
0.435
0.523
0.545
0.555
0.524
0.500
rubric
0.277
0.289
0.245
0.270
0.290
0.343
0.353
Key (exact key accuracy)
objective
0.043
0.058
0.099
0.123
0.126
0.112
–
rubric
0.045
0.047
0.048
0.040
0.045
0.045
–
Instrument. (contrastive)
objective
0.414
0.498
0.592
0.538
–
–
–
rubric
0.414
0.436
0.443
0.402
–
–
–
Appendix
Table 5: Values plotted in Figure 3 : the objective measure of each attribute on the evaluation prompts after the given number of DiffusionNFT updates, for the rubric-reward and objective-reward conditions (bold: peak within the plotted budget). Update 0 is the untrained model. On instrumentation, the rubric condition also raises the direct requested-instrument probability (mean classifier probability of the requested instruments) from 0.246 to 0.287 at its peak.
We introduce TuneJury, an open, instance-level pairwise reward model for text-to-music that predicts a music preference score from a text prompt and an audio clip. The released checkpoint is trained on publicly available human-preference labels covering arena-style (A vs. B) votes, metric-alignment preference pairs, crowdsourced pairwise comparisons, and expert aesthetic ratings. The predicted score margin between two clips is well calibrated on our held-out test split, supporting data filtering via a simple score threshold. TuneJury generalizes to both held-out test pairs and out-of-distribution benchmarks, remaining competitive with prior baselines on the latter. For generators released after training, we introduce anchor calibration, a post-hoc, per-system Bradley-Terry calibration that recovers agreement at substantially better data efficiency than from-scratch retraining. The same frozen reward drives consistent reward-axis gains across three downstream applications: inference-time best-of-N selection, DITTO-style latent optimization, and expert-iteration post-training. TuneJury is available at https://github.com/yonghyunk1m/TuneJury.
We describe our entry to the efficiency track of the Academic Text-to-Music (ATTM) Grand Challenge at ICME 2026. Beyond the challenge protocol's FAD-CLAP and CLAP score, we add a learned human-preference reward from TuneJury, a twin pairwise ranker trained over open music-preference datasets. The reward serves both as a training-time conditioning signal and as a sample-selection criterion. The pipeline combines five engineering decisions on a 120M-parameter FluxAudio-S backbone, four at training time and one at inference: (i) training-time reward conditioning that doubles as an inference-time CFG axis, (ii) a sweep over five score-conditioning architectures, where training and inference use different variants, (iii) expert iteration on the top decile, (iv) a short preference-tuning pass (CRPO) for audio-text alignment, and (v) inference post-processing via joint CFG, source separation, and loudness normalization. Per-stage decomposition on 100 Song Describer prompts shows training-time reward conditioning as a functional conditioning axis, expert iteration as the dominant contributor, the preference-tuning pass adding only noise-level gain, and the inference-time score scalar already saturated by the end of the chain.
Text-to-music generation has advanced rapidly, with modern autoregressive and diffusion-based models producing convincing music from natural-language prompts. However, much of this progress relies on large-scale training data and external pretraining, making it difficult to isolate which design choices remain effective when data and pretraining are controlled. We study this setting using a Diffusion Transformer backbone with lyric and timbre conditioning, adapted to an instrumental-only text-to-music task in which the auxiliary lyric and timbre branches receive only degenerate conditioning signals. Through controlled ablations, we find that models retrained without these branches score lower across AudioBox aesthetics, LLM-as-judge, and human MOS, and that reinvesting the saved parameters as additional DiT depth recovers only marginally. This suggests the auxiliary branches may act as training-time architectural anchors whose contribution goes beyond their explicit conditioning content. We validate the same model through comparisons with external instrumental baselines and through our submission to the ICME 2026 Academic Text-to-Music (ATTM) Grand Challenge, where our Performance submission ranked first under both the objective metrics and the subsequent organizer-administered MOS over 35 raters, attaining the highest overall MOS across all challenge submissions, while our Efficiency submission was a finalist that tied for second under the objective metrics.
Junyoung Koh
Department of Artificial Intelligence, Yonsei University MAAP KRAFTON Seoul, Republic of Korea