CASM: Context-Aware Semi-Markov Post-Processor for Beat Tracking
Authors: Zhanhong He, Hanyu Meng, Yaolong Ju
Organizations: The University of Western Australia, Perth, Australia · The University of New South Wales, Sydney, Australia · Great Bay University, Dongguan, China
In music beat tracking, model-predicted activations must be decoded into a discrete, musically coherent event sequence. Direct peak picking closely follows local evidence but can retain spurious peaks or miss weak beats. Dynamic Bayesian networks (DBNs), a widely used structured post-processor, improve sequence consistency under predefined global tempo, meter, and transition constraints, but their behavior can depend strongly on these settings. We introduce CASM, a context-aware semi-Markov decoder that instead conditions its temporal constraint on local activation evidence. CASM also accounts for ambiguity among competing periodic interpretations, including half- and double-tempo alternatives. Deterministic safeguards prevent implausible outputs and preserve beat-downbeat consistency. Applied to fixed activations from three neural beat trackers (BeatThis, MSCNN, and TCN), CASM improves temporal continuity while preserving event-level F1 across the GTZAN and SMC datasets, without backbone retraining or dataset-specific retuning. Further analysis shows that CASM is less sensitive than the DBN baseline to the composition of the calibration data.
Figures & tables
Figure 1: CASM behavior on BeatThis activations. (a) CASM enforces stronger temporal regularization when periodicity is clear, but (b) closely follows direct peak-picking outputs when periodicity is ambiguous. w(ci) denotes the strength of temporal regularization.
Method
GTZAN (Beat / Downbeat)
SMC (Beat)
F1
CMLt
AMLt
F1
CMLt
AMLt
F1
CMLt
AMLt
BeatThis [ 13 ]
89.1
79.8
89.8
78.3
67.2
79.1
62.7
51.4
61.0
w. CASM
+ 0.3
+ 0.8
+ 1.2
+ 0.7
+ 3.6
+ 8.0
+ 0.3
+ 2.4
+ 3.6
w. DBN [ 21 ]
− 0.5
− 0.9
+ 0.2
− 1.5
− 0.6
+ 1.1
− 1.6
+ 0.4
+ 3.5
w. DBN [55,215]
− 1.0
+ 0.7
+ 1.3
− 0.9
+ 6.2
+ 8.7
− 5.2
− 3.9
+ 1.5
w. PLPDP [ 7 ]
+ 0.1
+ 0.3
+ 0.3
+ 0.1
+ 0.2
+ 0.2
− 0.9
− 3.5
− 0.3
Table 1: Post-processing comparison on fixed backbone activations. Backbone-only rows use direct peak picking. Post-processors use 30–300 BPM unless marked [55,215] , the default DBN range which optimized for GTZAN-like music.
Variant
Sec.
GTZAN (Beat)
SMC (Beat)
F1
CMLt
AMLt
F1
CMLt
AMLt
CASM (full)
–
+ 0.3
+ 0.8
+ 1.2
+ 0.3
+ 2.4
+ 3.6
55–215-BPM support
2.2
+ 0.1
+ 0.5
+ 0.6
+ 0.1
+ 0.4
+ 2.0
Fixed stiffness (no margin)
2.3
− 0.3
− 0.8
− 0.8
+ 0.1
+ 0.7
+ 0.9
Margin on weight only
2.4
− 0.1
− 0.2
− 0.2
− 0.5
+ 0.9
+ 2.1
Margin on tolerance only
2.4
− 0.3
− 0.6
− 0.6
− 0.5
+ 0.9
+ 2.2
Table 2: Ablation of CASM on BeatThis activations. Each row changes one design choice described in the subsection of Sec. 2 . Values are changes from the direct peak-picking in percentage points.
Method
GTZAN (Beat / Downbeat)
SMC (Beat)
F1
CMLt
AMLt
F1
CMLt
AMLt
F1
CMLt
AMLt
HN-MERT [ 27 ]
89.7
81.4
94.3
77.4
71.6
87.2
61.7
–
–
HN-MusicFM
89.2
80.9
93.7
79.8
73.2
89.5
60.6
–
–
BeatFM-MERT [ 26 ]
89.5
82.7
93.9
76.7
70.2
85.4
61.3
–
–
BeatFM-MusicFM
89.1
80.6
93.5
79.6
74.4
88.7
60.5
–
–
BeatMamba [ 28 ]
88.9
82.3
94.3
–
–
–
58.7
46.8
62.4
Table 3: Comparison with recent beat-tracking systems. BeatThis (direct) and CASM results were obtained from our experiments, while all other results are reported from their original papers.
Figure 2: Sensitivity to calibration data for CASM and DBN. Each point is one decoder configuration, calibrated on one of the 7/21/35 possible 1-/2-/4-fold subsets of SMC folds 1–7 (7F: all seven folds, a single configuration) and evaluated on SMC fold 0 and GTZAN. Because the subsets overlap, the boxes describe spread only. Dashed lines show the direct peak-picking performance as the reference.
Current neural networks for beat tracking generate invalid outputs, such as consecutive downbeats and erratic tempo changes, even when these are not present in the training data. Heavy post-processing techniques can alleviate these problems, but the original cause of this inconsistent behaviour remains unknown. We hypothesise that it stems from inadequate modelling of multiple plausible output beat grids, resulting in an invalid mixture of competing interpretations. We propose a masked diffusion approach that properly models multiple outputs and enables the model to build coherent predictions through iterative inference. We devise three modifications to standard masked diffusion that enable its application to beat tracking: independent masking of beats and downbeats during training and inference, a balanced masking scheduler for inference, and peak-picking across inference steps. Our approach reduces erratic behaviours and improves beat-tracking performance.
Francesco Foscarin, Filip Korzeniowski, Richard Vogl
Over the past two decades, the task of musical beat tracking has transitioned from heuristic onset detection algorithms to highly capable deep neural networks (DNN). Although DNN-based beat tracking models achieve near-perfect performance on mainstream, percussive datasets, the SMC dataset has stubbornly yielded low F-measure scores. By testing how well state-of-the-art models detect beats on individual tracks in the SMC dataset, we identify three distinct failure modes: octave errors, continuity errors, and complete tracking failure where all metrics fall below 0.3. We reveal that state-of-the-art models tend to generate "confident-but-wrong" activations. Furthermore, we show that the standard DBN's default minimum tempo of 55 BPM prevents it from inferring the correct tempo for 21% of SMC tracks, forcing double-tempo predictions on slow music. By exposing such fundamental oversights, we provide concrete directions for improving beat and downbeat detection, specifically emphasizing training data diversification and multi-hypothesis tempo estimation.
Music creation is fundamentally a process of revision. Yet symbolic music generation remains dominated by paradigms that produce complete sequences from scratch, with limited support for selective modification. Edit-based methods have proven effective for text transformation tasks, but remain largely unexplored for symbolic music. We trace this absence to the representational level: conventional event-based music encodings lack the structural properties required by explicit music editing. In contrast, the BEAT encoding, a beat-grid-anchored representation originally designed for autoregressive generation, possesses structural properties amenable to editing. We propose BeatEdit, the first framework for symbolic music generation based on explicit edit operations, recasting generation as producing new content by editing a draft rather than synthesizing from scratch. BeatEdit comprises three complementary mechanisms along an axis of increasing edit density: per-token sequence tagging for error correction, iterative refinement for accompaniment editing, and tag-then-fill for segment completion. All these mechanisms share a single encoding and pre-trained backbone, achieving higher precision and perceptual quality than autoregressive and diffusion methods across all three tasks, while remaining efficient, with single-pass inference completing in under 100 ms. Cross-encoding evaluation further reveals that encoding design substantially influences editing effectiveness, with notable encoding-method interaction effects. Code is available at https://github.com/Haoyu-Gu/BeatEdit-code
Haoyu Gu, Lekai Qian, Haowu Zhou +2
School of Future Technology South China University of Technology Guangzhou, China · School of Intelligence Science and Technology Nanjing University Suzhou, China