Current music generative models can produce high-quality music, but does this ability imply that they ``understand'' the musical qualities of their outputs, and is that understanding aligned with human evaluation? Previous attempts to use the likelihood of a generative model to evaluate music, an approach commonly used in text, have proven unsuccessful, leading researchers to rely on standalone supervised music evaluation models. In this paper, we answer this question affirmatively: we show that a model's intrinsic signals---derived from its hidden representations and predictions---are strongly correlated with human ratings. In particular, we study MusicGen and consider three types of features: (1) prediction loss, (2) prediction entropy, and (3) concepts extracted from the model using a sparse autoencoder (SAE). Using these features, we train a lightweight prediction model to estimate subjective ratings. We evaluate these features both individually and in combination. We hypothesize that these signals parallel the listening process: the temporal and frequency-domain structure of loss and entropy reflects listeners' expectation and surprise, while gradient directions in SAE latent space predict perceived quality. Experiments on five human-evaluation benchmarks spanning continuous ratings and pairwise preferences confirm this hypothesis, with SAE latents carrying most of the predictive signal.
Figures & tables
Figure 1 : Hybrid prediction model: one 1D-CNN encoder per intrinsic signal, late concatenation (Eq. 3 ), and a shared MLP head producing y^∈[1,5] .
Figure 2 : Experimental flowchart: human-evaluation data, per-clip loss / entropy / SAE features, and the hybrid network with its ablations, evaluated against human ratings.
Group
Model
MusicEval
SongEval
AIME
MusicPref
Music Arena
All benchmarks
r
ρ
r
ρ
r
ρ
r
ρ
r
ρ
r
ρ
Baseline
Aesthetics-CE
0.62
0.64
0.57
0.57
0.41
0.38
0.04
0.03
0.24
0.26
0.20
0.21
Aesthetics-CU
0.62
0.64
0.63
0.70
0.48
0.45
0.04
0.04
0.22
0.24
0.18
0.19
Aesthetics-PC
0.08
0.04
0.32
0.27
0.01
0.01
0.01
0.00
0.15
0.23
0.07
0.09
Aesthetics-PQ
0.56
0.59
0.61
0.65
0.42
0.40
0.03
0.03
0.33
0.40
0.20
0.23
Mean Loss
-0.02
-0.07
0.24
0.11
-0.10
-0.09
-0.30
-0.41
-0.08
-0.07
-0.08
-0.13
Table 1 : Pearson r and Spearman ρ vs. human evaluation; trained cells are MusicGen-small / large. Trained rows: held-out test split per benchmark; Audiobox rows: zero-shot on the full set; All benchmarks : union of the five held-out splits. Per size, bold denotes the best score in each column, and underlining denotes scores tied with it (bootstrap, p≥0.05 [ 30 ] ).
Figure 3 : Example clip ( y=4.9 ). (a) 100-token window with the densest co-occurrence of CW and UW tokens. (b) Full-length (ℓt,Ht) plane for the same clip.
per-clip descriptor
r
ρ
π~CC (confident-correct)
+0.03
+0.02
π~UC (uncertain-correct)
−0.12†
−0.12†
π~CW (overconfident-wrong)
−0.18‡
−0.19‡
π~UW (uncertain-wrong)
+0.29✠
+0.32✠
π~UW−π~CW
+0.25✠
+0.27✠
ℓ−H (excess surprise)
+0.30✠
+0.35✠
Table 2 : Per-clip loss–entropy configurations vs. human ratings (Pearson r , Spearman ρ ). †p<0.05 ; ‡p<0.01 ; ✠p<0.001 ; unmarked: p≥0.05 .
Figure 4 : Spectral correlates of rating: Pearson r between per-band power of ℓt,Ht and y across log-spaced rate bands. Left: 64 fine bands. Right: five-band roll-up.
rank
ρk
VLM caption
producer concept
good-aligned ( ρk>0 )
g-1
+0.54
classical, string
string classical
g-2
+0.49
folk, upbeat
folk guitar texture
g-5
+0.45
folk, calm
choir + plucked folk
g-7
+0.44
folk, melancholic
flute texture
g-9
+0.43
ambient, calm
ambient pads + vocal tex.
Table 3 : Representative SAE latents ranked by rating-aligned attribution ρk , each annotated by Gemini- 2.5 -Flash and by a professional music producer. Thin red/blue bars in the ρk column visualize ∣ρk∣ .