Current music generative models can produce high-quality music, but does this ability imply that they ``understand'' the musical qualities of their outputs, and is that understanding aligned with human evaluation? Previous attempts to use the likelihood of a generative model to evaluate music, an approach commonly used in text, have proven unsuccessful, leading researchers to rely on standalone supervised music evaluation models. In this paper, we answer this question affirmatively: we show that a model's intrinsic signals---derived from its hidden representations and predictions---are strongly correlated with human ratings. In particular, we study MusicGen and consider three types of features: (1) prediction loss, (2) prediction entropy, and (3) concepts extracted from the model using a sparse autoencoder (SAE). Using these features, we train a lightweight prediction model to estimate subjective ratings. We evaluate these features both individually and in combination. We hypothesize that these signals parallel the listening process: the temporal and frequency-domain structure of loss and entropy reflects listeners' expectation and surprise, while gradient directions in SAE latent space predict perceived quality. Experiments on five human-evaluation benchmarks spanning continuous ratings and pairwise preferences confirm this hypothesis, with SAE latents carrying most of the predictive signal.
Figures & tables
Figure 1 : Hybrid prediction model: one 1D-CNN encoder per intrinsic signal, late concatenation (Eq. 3 ), and a shared MLP head producing y^∈[1,5] .
Figure 2 : Experimental flowchart: human-evaluation data, per-clip loss / entropy / SAE features, and the hybrid network with its ablations, evaluated against human ratings.
Group
Model
MusicEval
SongEval
AIME
MusicPref
Music Arena
All benchmarks
r
ρ
r
ρ
r
ρ
r
ρ
r
ρ
r
ρ
Baseline
Aesthetics-CE
0.62
0.64
0.57
0.57
0.41
0.38
0.04
0.03
0.24
0.26
0.20
0.21
Aesthetics-CU
0.62
0.64
0.63
0.70
0.48
0.45
0.04
0.04
0.22
0.24
0.18
0.19
Aesthetics-PC
0.08
0.04
0.32
0.27
0.01
0.01
0.01
0.00
0.15
0.23
0.07
0.09
Aesthetics-PQ
0.56
0.59
0.61
0.65
0.42
0.40
0.03
0.03
0.33
0.40
0.20
0.23
Mean Loss
-0.02
-0.07
0.24
0.11
-0.10
-0.09
-0.30
-0.41
-0.08
-0.07
-0.08
-0.13
Table 1 : Pearson r and Spearman ρ vs. human evaluation; trained cells are MusicGen-small / large. Trained rows: held-out test split per benchmark; Audiobox rows: zero-shot on the full set; All benchmarks : union of the five held-out splits. Per size, bold denotes the best score in each column, and underlining denotes scores tied with it (bootstrap, p≥0.05 [ 30 ] ).
Figure 3 : Example clip ( y=4.9 ). (a) 100-token window with the densest co-occurrence of CW and UW tokens. (b) Full-length (ℓt,Ht) plane for the same clip.
per-clip descriptor
r
ρ
π~CC (confident-correct)
+0.03
+0.02
π~UC (uncertain-correct)
−0.12†
−0.12†
π~CW (overconfident-wrong)
−0.18‡
−0.19‡
π~UW (uncertain-wrong)
+0.29✠
+0.32✠
π~UW−π~CW
+0.25✠
+0.27✠
ℓ−H (excess surprise)
+0.30✠
+0.35✠
Table 2 : Per-clip loss–entropy configurations vs. human ratings (Pearson r , Spearman ρ ). †p<0.05 ; ‡p<0.01 ; ✠p<0.001 ; unmarked: p≥0.05 .
Figure 4 : Spectral correlates of rating: Pearson r between per-band power of ℓt,Ht and y across log-spaced rate bands. Left: 64 fine bands. Right: five-band roll-up.
rank
ρk
VLM caption
producer concept
good-aligned ( ρk>0 )
g-1
+0.54
classical, string
string classical
g-2
+0.49
folk, upbeat
folk guitar texture
g-5
+0.45
folk, calm
choir + plucked folk
g-7
+0.44
folk, melancholic
flute texture
g-9
+0.43
ambient, calm
ambient pads + vocal tex.
Table 3 : Representative SAE latents ranked by rating-aligned attribution ρk , each annotated by Gemini- 2.5 -Flash and by a professional music producer. Thin red/blue bars in the ρk column visualize ∣ρk∣ .
Music popularity prediction has attracted growing research interest, with relevance to artists, platforms, and recommendation systems. However, the explosive rise of AI-generated music platforms has created an entirely new and largely unexplored landscape, where a surge of songs is produced and consumed daily without the traditional markers of artist reputation or label backing. Key, yet unexplored in this pursuit is aesthetic quality. We propose APEX, the first large-scale multi-task learning framework for AI-generated music, trained on over 211k songs (10k hours of audio) from Suno and Udio, that jointly predicts engagement-based popularity signals - streams and likes scores - alongside five perceptual aesthetic quality dimensions from frozen audio embeddings extracted from MERT, a self-supervised music understanding model. Aesthetic quality and popularity capture complementary aspects of music that together prove valuable: in an out-of-distribution evaluation on the Music Arena dataset, comprising pairwise human preference battles across eleven generative music systems unseen during training, including aesthetic features consistently improves preference prediction, demonstrating strong generalisation of the learned representations across generative architectures.
Jaavid Aktar Husain, Dorien Herremans
AMAAI Lab, Singapore University of Technology and Design
Long-form song generation models continue to improve in duration, structural coherence, and acoustic complexity, increasing the need for reliable aesthetic rewards aligned with human preferences. However, reward models for complete songs remain limited, and existing evaluators typically predict scores in a single forward pass without readable explanations. To this end, we introduce MuseCritic, a semi-scalar reward model that generates a natural-language critique covering five aesthetic dimensions and uses it as an intermediate representation to predict continuous reward scores. MuseCritic follows a two-stage training pipeline: a teacher model first provides high-quality critiques for supervised fine-tuning, then the fine-tuned model generates its own critiques for reward learning, mitigating training-inference distribution shift. On an in-domain test set of 200 SongEval songs, MuseCritic reduces macro-averaged mean squared error from 0.2875 to 0.2316 and improves macro-averaged LCC, SRCC, and Kendall's tau to 0.9068, 0.8838, and 0.7178, respectively. On the out-of-domain Music Arena benchmark with 733 preference pairs, it achieves 71.35% accuracy and remains competitive with strong music-specific reward models. Using MuseCritic with GRPO also improves Muse-0.6B on all nine aesthetic metrics from SongEval and Audiobox Aesthetics. These results show that critique-conditioned reward modeling reduces scoring error and provides an effective optimization signal for song generation. The project repository is available at https://github.com/WuqnEl/MuseCritic.
Music aesthetics scoring plays a critical role in applications such as dataset curation, generative model evaluation, and reward modeling for music generation. Recent approaches rely on deep neural networks trained on human-annotated ratings, but these models may exploit spurious correlations rather than capturing perceptually meaningful aesthetics. In this work, we identify a previously underexplored failure mode in music evaluation models: genre-induced shortcut learning. Through a systematic analysis of SongEval, we show that biases in training data lead to strong correlations between genre-related features and predicted scores, causing the model to use them as a proxy for aesthetics. This results in systematic overestimation of pop music and undervaluation of high-quality samples from other genres, leading to predictions that are inconsistent with human preferences. To address this issue, we propose a training objective that jointly reweights hard samples and regularizes group-level performance, encouraging the model to learn genre-invariant representations of musicality. Experimental results demonstrate that our method reduces genre-dependent bias and improves alignment with human preferences, as reflected by gains in both cross-genre and within-genre preference alignment.
Yizhou Zhang, Wangjin Zhou, Yi Zhao +3
Graduate School of Informatics, Kyoto University, Japan · WXG, Tencent, China