MIRA: A Musical Intent Refinement Agent for Aligning Text-to-Music Generation with User Intent
Authors: Zekai Liu, Zhilin Wang, Xuzheng He, Yu Cheng, Yang Yang
Organizations: Shandong University · University of Science and Technology of China · Central Conservatory of Music · Kunlun Tech Co. Ltd. · Shanghai Jiao Tong University
Text-to-music systems produce increasingly convincing audio, yet evaluation reveals little about whether the result matches user intent. A global text-audio relevance score can overlook the implicit intent in underspecified prompts and mask failures in specific requirements, such as instrumentation, structure, rhythm, or mood progression. To bridge this gap, we formulate text-to-music intent alignment as satisfying a per-request rubric of independently verifiable items covering both a request's explicit requirements and its implied musical intent. Scoring items individually makes evaluation diagnostic by intent source and musical dimension, rather than a single opaque score. We instantiate this as MuRA-Bench, a benchmark of real-world platform requests curated by music experts. We further propose MIRA (Musical Intent Refinement Agent), a test-time agent that first grounds a request's intent into rubrics, then searches over prompt revisions for a black-box generator under a bounded budget, iteratively generating music, verifying it against the rubrics, and using this feedback to guide a trajectory-aware tree search. Experiments across open-source and commercial backends show that MIRA improves intent alignment, enabling an open-source generator to achieve performance comparable to representative commercial systems (e.g. Suno and Mureka). Project page: https://mirareview.github.io/.
Figures & tables
Figure 1: Overview of MuRA-Bench construction and evaluation. Expert-revised, request-specific rubrics support fine-grained assessment with MusicFlamingo. Gold rubrics are used exclusively for evaluation and are not exposed to the generation system.
Figure 2: Overview of MIRA and its verifier-guided refinement process.
Item-level
Clip-level
Pair
Evaluator
SRCC
τb
LCC
SRCC
τb
Acc. (%)
MusicFlamingo
.690
.532
.817
.823
.646
79.59
Qwen3-Omni-30B-A3B
.509
.375
.632
.631
.480
74.15
Qwen2.5-Omni-7B
.469
.344
.599
.643
.465
72.79
MOSS-Music-8B
.502
.371
.591
.593
.429
62.59
AnyAudio-Judge-7B
.462
.335
.561
.556
.393
68.03
Table 1: Agreement with expert ratings on 100 clips from 25 requests. The first five evaluators use AQA. MF denotes MusicFlamingo.
Intent alignment
Musical dimensions
Backend / Setting
Overall
D-Macro
Base
Comp.
Style
I/V
Mood
Rhythm
H/M
S/E
P/T
ACE-Step v1.5 Turbo
Direct
67.6
67.2
69.6
66.1
68.1
62.4
71.9
71.0
64.1
69.4
63.3
MIRA ( B=3 )
74.2 ↑ 6.6
74.1 ↑ 6.9
77.1 ↑ 7.5
71.1 ↑ 5.0
74.6 ↑ 6.5
69.5 ↑ 7.1
81.2 ↑ 9.3
75.4 ↑ 4.4
71.4 ↑ 7.3
72.8 ↑ 3.4
73.7 ↑ 10.4
MIRA ( B=9 )
76.3 ↑ 8.7
75.6 ↑ 8.4
79.1 ↑ 9.5
73.4 ↑ 7.3
75.4 ↑ 7.3
69.1 ↑ 6.7
83.0 ↑ 11.1
77.2 ↑ 6.2
72.8 ↑ 8.7
74.6 ↑ 5.2
77.0 ↑ 13.7
SongGeneration2 Large
Table 2: Results on the full MuRA-Bench benchmark. All scores are multiplied by 100; higher is better. Superscripts in MIRA rows indicate absolute changes from the same backend’s direct-prompt baseline on this scale. Bold marks the best score within each refinement backend or among the commercial models.
Figure 3: Blinded expert ratings of overall, explicit, and inferred intent on MuRA-Bench. Bars show mean ratings on a 1–5 scale, with 95% confidence intervals.
ACE-Step v1.5 Turbo
YuE2-3B
Setting
B
Overall
D-Macro
Base
Comp.
Overall
D-Macro
Base
Comp.
Direct
1
0.759
0.756
0.803
0.710
0.816
0.811
0.868
0.752
Best-of-9
9
0.800
0.795
0.874
0.719
0.848
0.850
0.910
0.775
Tools + Best-of-9
9
0.815
0.807
0.871
0.748
0.860
0.852
0.911
0.800
Tools + Adaptive search
9
0.826
0.814
0.862
0.781
0.867
0.863
0.913
0.808
Full MIRA
9
0.836
0.826
0.880
0.782
0.872
0.868
0.911
0.821
Table 3: Component analysis on a fixed subset of 20 MuRA-Bench requests. Direct generation uses B=1 ; all remaining configurations use B=9 . Tool grounding is first combined with best-of-nine sampling, followed by adaptive search without memory and then tree-structured memory. Bold marks the best score within each backend.
Figure 4: Best-of-9 versus MIRA with B=9 . Scores are multiplied by 100; the vertical axis starts at 75.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Rubric dimension
Items
Style
135
Instrumentation/vocal
303
Mood
218
Rhythm
257
Harmony/melody
114
Structure/energy
199
Appendix
Table 4: Distribution of the 1,566 rubric items across seven musical dimensions.
Source
Requirement
Dimension
Base
upbeat combat music for a 2D role-playing game
Style
Base
open-plains biome atmosphere
Mood
Base
flute parts
Instrumentation/vocal
Base
wind textures
Production/texture
Base
African drum rhythms
Rhythm
Base
energizing battle-ready mood
Mood
Appendix
Table 5: A complete expert-revised rubric for RPG battle music, comprising six base and six completion requirements.
20-point
5-point
Interpretation
1–4
1
Not fulfilled: the required property is absent or clearly contradicted.
5–8
2
Slightly fulfilled: only weak evidence is audible, with most of the requirement unmet.
9–12
3
Partially fulfilled: recognizable evidence is present, but substantial shortcomings remain.
13–16
4
Mostly fulfilled: the main requirement is realized, with minor shortcomings.
17–20
5
Fully fulfilled: the requirement is clearly and sufficiently realized, without an evident deviation.
Appendix
Table 6: Scoring anchors for item-level and system-level expert assessments.
Setting
Value or rule
Generation budget B
3 or 9, including the initial candidate
Low-score feedback threshold η
0.60
Memory change tolerance ϵ
0.03
Online ranking score
Mean item satisfaction over the fixed online rubric
Complementary frontier criterion
Mean of the lowest-scoring quarter of items
Final selection
Highest-scoring successfully visited candidate
Appendix
Table 7: MIRA search settings and generation-budget conventions.
Intent alignment
Musical dimensions
Backend / Setting
Overall
D-Macro
Base
Comp.
Style
I/V
Mood
Rhythm
H/M
S/E
P/T
Evaluator: MOSS-Music
ACE-Step v1.5 Turbo
Direct
64.7
65.6
66.5
63.8
69.0
55.8
68.9
67.5
69.1
69.7
59.0
MIRA ( B=3 )
71.2 ↑ 6.5
71.5 ↑ 5.9
73.1 ↑ 6.7
68.1 ↑ 4.3
73.3 ↑ 4.3
66.0 ↑ 10.2
80.1 ↑ 11.2
70.5 ↑ 3.0
70.9 ↑ 1.8
70.6 ↑ 0.9
68.8 ↑ 9.8
MIRA ( B=9 )
71.9 ↑ 7.1
72.7 ↑ 7.1
73.1 ↑ 6.6
69.8 ↑ 6.0
74.5 ↑ 5.5
59.9 ↑ 4.1
80.4 ↑ 11.4
72.2 ↑ 4.7
74.8 ↑ 5.7
73.8 ↑ 4.2
73.2 ↑ 14.2
Appendix
Table 8: Cross-evaluator assessment on MuRA-Bench. MOSS-Music and Qwen3-Omni score outputs from direct generation and MusicFlamingo-guided MIRA. Commercial direct-prompting baselines are evaluated with MOSS-Music and Qwen3-Omni. Scores are multiplied by 100; higher is better. Superscripts indicate differences from the reported Direct score within each backend–evaluator group. Bold marks the best reported score within each refinement group or among the commercial baselines, before rounding.
Aug 10, 2026·Shuyu Li, Kejun Zhang, Jiahe Lei +5Text-To-Music
College of Artificial Intelligence, Zhejiang University · 3Innovation Center of Yangtze River Delta, Zhejiang University · 4The Chinese University of Hong Kong +4
Oct 2, 2026·Ping Wang, Guang Yang, Shao-Rong Su +3Text-To-Music
Paul G. Allen School of Computer Science & Engineering, University of Washington · Department of Electrical and Computer Engineering, University of Washington · Allen Institute for AI