cs.SDOct 2, 2026

Rubric-Based Optimization for Text-to-Music Generation

Authors: Ping Wang, Guang Yang, Shao-Rong Su, Junkai Wu, Pang Wei Koh, Noah A. Smith

Organizations: Paul G. Allen School of Computer Science & Engineering, University of Washington · Department of Electrical and Computer Engineering, University of Washington · Allen Institute for AI

Abstract

Post-training text-to-music generation requires reward signals that capture multiple aspects of musical quality beyond what any single automatic metric can measure. We study structured, rubric-based rewards from pretrained audio-language models (ALMs) as training signals for both autoregressive and diffusion-based music generators. An ALM scores each generated clip against the rubric; we rank candidates generated for the same text prompt by their scores and convert these rankings into preference pairs for DPO on both MusicGen-small and ACE-Step v1, and additionally use the rubric scores directly as scalar rewards for DiffusionNFT on ACE-Step v1. On MusicCaps, rubric-based optimization improves CLAP, SongEval, and Audiobox-Aesthetics simultaneously, with the strongest gains obtained by DiffusionNFT on ACE-Step. By contrast, on MusicGen-small, building preferences from any one of these automatic evaluators produces clear cross-metric trade-offs: the targeted evaluator improves while other independent evaluators deteriorate. We further study tempo, key, and instrumentation, where precise objective rewards are available. Directly optimizing these specialized rewards reliably improves the target attributes, whereas ALM rubrics provide only partial transfer for tempo and instrumentation and no measurable improvement for key. Together, these results suggest a practical division of labor: ALM rubrics are effective for broad perceptual qualities that are difficult to formalize, while specialized objective rewards remain preferable when reliable measurements are available.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TuneJury: An Open Metric for Improving Music Generation Preference Alignment

    Jun 15, 2026Yonghyun Kim, Junwon Lee, Haiwen Xia +5Text-To-MusicSong Generation

  2. Improving Text-to-Music Generation with Human Preference Rewards

    Jun 19, 2026Yonghyun Kim, Junwon Lee, Haiwen Xia +2Text-To-MusicPreference Datasets

  3. Instrumental Text-to-Music Generation with Auxiliary Conditioning Branches

    May 20, 2026Junyoung KohText-To-Music