MIRA: A Musical Intent Refinement Agent for Aligning Text-to-Music Generation with User Intent
Authors: Zekai Liu, Zhilin Wang, Xuzheng He, Yu Cheng, Yang Yang
Organizations: Shandong University · University of Science and Technology of China · Central Conservatory of Music · Kunlun Tech Co. Ltd. · Shanghai Jiao Tong University
Text-to-music systems produce increasingly convincing audio, yet evaluation reveals little about whether the result matches user intent. A global text-audio relevance score can overlook the implicit intent in underspecified prompts and mask failures in specific requirements, such as instrumentation, structure, rhythm, or mood progression. To bridge this gap, we formulate text-to-music intent alignment as satisfying a per-request rubric of independently verifiable items covering both a request's explicit requirements and its implied musical intent. Scoring items individually makes evaluation diagnostic by intent source and musical dimension, rather than a single opaque score. We instantiate this as MuRA-Bench, a benchmark of real-world platform requests curated by music experts. We further propose MIRA (Musical Intent Refinement Agent), a test-time agent that first grounds a request's intent into rubrics, then searches over prompt revisions for a black-box generator under a bounded budget, iteratively generating music, verifying it against the rubrics, and using this feedback to guide a trajectory-aware tree search. Experiments across open-source and commercial backends show that MIRA improves intent alignment, enabling an open-source generator to achieve performance comparable to representative commercial systems (e.g. Suno and Mureka). Project page: https://mirareview.github.io/.
Figures & tables
Figure 1: Overview of MuRA-Bench construction and evaluation. Expert-revised, request-specific rubrics support fine-grained assessment with MusicFlamingo. Gold rubrics are used exclusively for evaluation and are not exposed to the generation system.
Figure 2: Overview of MIRA and its verifier-guided refinement process.
Item-level
Clip-level
Pair
Evaluator
SRCC
τb
LCC
SRCC
τb
Acc. (%)
MusicFlamingo
.690
.532
.817
.823
.646
79.59
Qwen3-Omni-30B-A3B
.509
.375
.632
.631
.480
74.15
Qwen2.5-Omni-7B
.469
.344
.599
.643
.465
72.79
MOSS-Music-8B
.502
.371
.591
.593
.429
62.59
AnyAudio-Judge-7B
.462
.335
.561
.556
.393
68.03
Table 1: Agreement with expert ratings on 100 clips from 25 requests. The first five evaluators use AQA. MF denotes MusicFlamingo.
Intent alignment
Musical dimensions
Backend / Setting
Overall
D-Macro
Base
Comp.
Style
I/V
Mood
Rhythm
H/M
S/E
P/T
ACE-Step v1.5 Turbo
Direct
67.6
67.2
69.6
66.1
68.1
62.4
71.9
71.0
64.1
69.4
63.3
MIRA ( B=3 )
74.2 ↑ 6.6
74.1 ↑ 6.9
77.1 ↑ 7.5
71.1 ↑ 5.0
74.6 ↑ 6.5
69.5 ↑ 7.1
81.2 ↑ 9.3
75.4 ↑ 4.4
71.4 ↑ 7.3
72.8 ↑ 3.4
73.7 ↑ 10.4
MIRA ( B=9 )
76.3 ↑ 8.7
75.6 ↑ 8.4
79.1 ↑ 9.5
73.4 ↑ 7.3
75.4 ↑ 7.3
69.1 ↑ 6.7
83.0 ↑ 11.1
77.2 ↑ 6.2
72.8 ↑ 8.7
74.6 ↑ 5.2
77.0 ↑ 13.7
SongGeneration2 Large
Table 2: Results on the full MuRA-Bench benchmark. All scores are multiplied by 100; higher is better. Superscripts in MIRA rows indicate absolute changes from the same backend’s direct-prompt baseline on this scale. Bold marks the best score within each refinement backend or among the commercial models.
Figure 3: Blinded expert ratings of overall, explicit, and inferred intent on MuRA-Bench. Bars show mean ratings on a 1–5 scale, with 95% confidence intervals.
ACE-Step v1.5 Turbo
YuE2-3B
Setting
B
Overall
D-Macro
Base
Comp.
Overall
D-Macro
Base
Comp.
Direct
1
0.759
0.756
0.803
0.710
0.816
0.811
0.868
0.752
Best-of-9
9
0.800
0.795
0.874
0.719
0.848
0.850
0.910
0.775
Tools + Best-of-9
9
0.815
0.807
0.871
0.748
0.860
0.852
0.911
0.800
Tools + Adaptive search
9
0.826
0.814
0.862
0.781
0.867
0.863
0.913
0.808
Full MIRA
9
0.836
0.826
0.880
0.782
0.872
0.868
0.911
0.821
Table 3: Component analysis on a fixed subset of 20 MuRA-Bench requests. Direct generation uses B=1 ; all remaining configurations use B=9 . Tool grounding is first combined with best-of-nine sampling, followed by adaptive search without memory and then tree-structured memory. Bold marks the best score within each backend.
Figure 4: Best-of-9 versus MIRA with B=9 . Scores are multiplied by 100; the vertical axis starts at 75.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Rubric dimension
Items
Style
135
Instrumentation/vocal
303
Mood
218
Rhythm
257
Harmony/melody
114
Structure/energy
199
Appendix
Table 4: Distribution of the 1,566 rubric items across seven musical dimensions.
Source
Requirement
Dimension
Base
upbeat combat music for a 2D role-playing game
Style
Base
open-plains biome atmosphere
Mood
Base
flute parts
Instrumentation/vocal
Base
wind textures
Production/texture
Base
African drum rhythms
Rhythm
Base
energizing battle-ready mood
Mood
Appendix
Table 5: A complete expert-revised rubric for RPG battle music, comprising six base and six completion requirements.
20-point
5-point
Interpretation
1–4
1
Not fulfilled: the required property is absent or clearly contradicted.
5–8
2
Slightly fulfilled: only weak evidence is audible, with most of the requirement unmet.
9–12
3
Partially fulfilled: recognizable evidence is present, but substantial shortcomings remain.
13–16
4
Mostly fulfilled: the main requirement is realized, with minor shortcomings.
17–20
5
Fully fulfilled: the requirement is clearly and sufficiently realized, without an evident deviation.
Appendix
Table 6: Scoring anchors for item-level and system-level expert assessments.
Setting
Value or rule
Generation budget B
3 or 9, including the initial candidate
Low-score feedback threshold η
0.60
Memory change tolerance ϵ
0.03
Online ranking score
Mean item satisfaction over the fixed online rubric
Complementary frontier criterion
Mean of the lowest-scoring quarter of items
Final selection
Highest-scoring successfully visited candidate
Appendix
Table 7: MIRA search settings and generation-budget conventions.
Intent alignment
Musical dimensions
Backend / Setting
Overall
D-Macro
Base
Comp.
Style
I/V
Mood
Rhythm
H/M
S/E
P/T
Evaluator: MOSS-Music
ACE-Step v1.5 Turbo
Direct
64.7
65.6
66.5
63.8
69.0
55.8
68.9
67.5
69.1
69.7
59.0
MIRA ( B=3 )
71.2 ↑ 6.5
71.5 ↑ 5.9
73.1 ↑ 6.7
68.1 ↑ 4.3
73.3 ↑ 4.3
66.0 ↑ 10.2
80.1 ↑ 11.2
70.5 ↑ 3.0
70.9 ↑ 1.8
70.6 ↑ 0.9
68.8 ↑ 9.8
MIRA ( B=9 )
71.9 ↑ 7.1
72.7 ↑ 7.1
73.1 ↑ 6.6
69.8 ↑ 6.0
74.5 ↑ 5.5
59.9 ↑ 4.1
80.4 ↑ 11.4
72.2 ↑ 4.7
74.8 ↑ 5.7
73.8 ↑ 4.2
73.2 ↑ 14.2
Appendix
Table 8: Cross-evaluator assessment on MuRA-Bench. MOSS-Music and Qwen3-Omni score outputs from direct generation and MusicFlamingo-guided MIRA. Commercial direct-prompting baselines are evaluated with MOSS-Music and Qwen3-Omni. Scores are multiplied by 100; higher is better. Superscripts indicate differences from the reported Direct score within each backend–evaluator group. Bold marks the best reported score within each refinement group or among the commercial baselines, before rounding.
We introduce TuneJury, an open, instance-level pairwise reward model for text-to-music that predicts a music preference score from a text prompt and an audio clip. The released checkpoint is trained on publicly available human-preference labels covering arena-style (A vs. B) votes, metric-alignment preference pairs, crowdsourced pairwise comparisons, and expert aesthetic ratings. The predicted score margin between two clips is well calibrated on our held-out test split, supporting data filtering via a simple score threshold. TuneJury generalizes to both held-out test pairs and out-of-distribution benchmarks, remaining competitive with prior baselines on the latter. For generators released after training, we introduce anchor calibration, a post-hoc, per-system Bradley-Terry calibration that recovers agreement at substantially better data efficiency than from-scratch retraining. The same frozen reward drives consistent reward-axis gains across three downstream applications: inference-time best-of-N selection, DITTO-style latent optimization, and expert-iteration post-training. TuneJury is available at https://github.com/yonghyunk1m/TuneJury.
Text-to-music generation has advanced rapidly, but current systems still rely primarily on global text prompts, leaving the structural organization of generated music implicit and difficult to inspect, control, or revise before audio generation. To address this issue, we introduce MusicLayout, an explicit intermediate representation for controlling musical structure in text-to-music generation. MusicLayout describes a musical piece as a time-aligned layout of sections, textures, repetitions, variations, and instrument-level arrangements, serving as an interpretable planning layer between textual intent and the generated music. We integrate MusicLayout into a text-to-music framework built on a unified autoregressive formulation, where the model first generates a MusicLayout representation and subsequently predicts audio tokens conditioned on this representation within a single sequence. The resulting MusicLayout can be inspected and modified prior to audio generation, providing a mechanism for layout-level structural control. We evaluate MusicLayout through layout-conditioned generation, layout manipulation experiments, and matched-data ablations, providing evidence that explicit layout planning can improve long-range structural organization and support layout-level control.
Shuyu Li, Kejun Zhang, Jiahe Lei +5
College of Artificial Intelligence, Zhejiang University · 3Innovation Center of Yangtze River Delta, Zhejiang University · 4The Chinese University of Hong Kong +4
Post-training text-to-music generation requires reward signals that capture multiple aspects of musical quality beyond what any single automatic metric can measure. We study structured, rubric-based rewards from pretrained audio-language models (ALMs) as training signals for both autoregressive and diffusion-based music generators. An ALM scores each generated clip against the rubric; we rank candidates generated for the same text prompt by their scores and convert these rankings into preference pairs for DPO on both MusicGen-small and ACE-Step v1, and additionally use the rubric scores directly as scalar rewards for DiffusionNFT on ACE-Step v1. On MusicCaps, rubric-based optimization improves CLAP, SongEval, and Audiobox-Aesthetics simultaneously, with the strongest gains obtained by DiffusionNFT on ACE-Step. By contrast, on MusicGen-small, building preferences from any one of these automatic evaluators produces clear cross-metric trade-offs: the targeted evaluator improves while other independent evaluators deteriorate. We further study tempo, key, and instrumentation, where precise objective rewards are available. Directly optimizing these specialized rewards reliably improves the target attributes, whereas ALM rubrics provide only partial transfer for tempo and instrumentation and no measurable improvement for key. Together, these results suggest a practical division of labor: ALM rubrics are effective for broad perceptual qualities that are difficult to formalize, while specialized objective rewards remain preferable when reliable measurements are available.
Ping Wang, Guang Yang, Shao-Rong Su +3
Paul G. Allen School of Computer Science & Engineering, University of Washington · Department of Electrical and Computer Engineering, University of Washington · Allen Institute for AI