CASM: Context-Aware Semi-Markov Post-Processor for Beat Tracking
Authors: Zhanhong He, Hanyu Meng, Yaolong Ju
Organizations: The University of Western Australia, Perth, Australia · The University of New South Wales, Sydney, Australia · Great Bay University, Dongguan, China
In music beat tracking, model-predicted activations must be decoded into a discrete, musically coherent event sequence. Direct peak picking closely follows local evidence but can retain spurious peaks or miss weak beats. Dynamic Bayesian networks (DBNs), a widely used structured post-processor, improve sequence consistency under predefined global tempo, meter, and transition constraints, but their behavior can depend strongly on these settings. We introduce CASM, a context-aware semi-Markov decoder that instead conditions its temporal constraint on local activation evidence. CASM also accounts for ambiguity among competing periodic interpretations, including half- and double-tempo alternatives. Deterministic safeguards prevent implausible outputs and preserve beat-downbeat consistency. Applied to fixed activations from three neural beat trackers (BeatThis, MSCNN, and TCN), CASM improves temporal continuity while preserving event-level F1 across the GTZAN and SMC datasets, without backbone retraining or dataset-specific retuning. Further analysis shows that CASM is less sensitive than the DBN baseline to the composition of the calibration data.
Figures & tables
Figure 1: CASM behavior on BeatThis activations. (a) CASM enforces stronger temporal regularization when periodicity is clear, but (b) closely follows direct peak-picking outputs when periodicity is ambiguous. w(ci) denotes the strength of temporal regularization.
Method
GTZAN (Beat / Downbeat)
SMC (Beat)
F1
CMLt
AMLt
F1
CMLt
AMLt
F1
CMLt
AMLt
BeatThis [ 13 ]
89.1
79.8
89.8
78.3
67.2
79.1
62.7
51.4
61.0
w. CASM
+ 0.3
+ 0.8
+ 1.2
+ 0.7
+ 3.6
+ 8.0
+ 0.3
+ 2.4
+ 3.6
w. DBN [ 21 ]
− 0.5
− 0.9
+ 0.2
− 1.5
− 0.6
+ 1.1
− 1.6
+ 0.4
+ 3.5
w. DBN [55,215]
− 1.0
+ 0.7
+ 1.3
− 0.9
+ 6.2
+ 8.7
− 5.2
− 3.9
+ 1.5
w. PLPDP [ 7 ]
+ 0.1
+ 0.3
+ 0.3
+ 0.1
+ 0.2
+ 0.2
− 0.9
− 3.5
− 0.3
Table 1: Post-processing comparison on fixed backbone activations. Backbone-only rows use direct peak picking. Post-processors use 30–300 BPM unless marked [55,215] , the default DBN range which optimized for GTZAN-like music.
Variant
Sec.
GTZAN (Beat)
SMC (Beat)
F1
CMLt
AMLt
F1
CMLt
AMLt
CASM (full)
–
+ 0.3
+ 0.8
+ 1.2
+ 0.3
+ 2.4
+ 3.6
55–215-BPM support
2.2
+ 0.1
+ 0.5
+ 0.6
+ 0.1
+ 0.4
+ 2.0
Fixed stiffness (no margin)
2.3
− 0.3
− 0.8
− 0.8
+ 0.1
+ 0.7
+ 0.9
Margin on weight only
2.4
− 0.1
− 0.2
− 0.2
− 0.5
+ 0.9
+ 2.1
Margin on tolerance only
2.4
− 0.3
− 0.6
− 0.6
− 0.5
+ 0.9
+ 2.2
Table 2: Ablation of CASM on BeatThis activations. Each row changes one design choice described in the subsection of Sec. 2 . Values are changes from the direct peak-picking in percentage points.
Method
GTZAN (Beat / Downbeat)
SMC (Beat)
F1
CMLt
AMLt
F1
CMLt
AMLt
F1
CMLt
AMLt
HN-MERT [ 27 ]
89.7
81.4
94.3
77.4
71.6
87.2
61.7
–
–
HN-MusicFM
89.2
80.9
93.7
79.8
73.2
89.5
60.6
–
–
BeatFM-MERT [ 26 ]
89.5
82.7
93.9
76.7
70.2
85.4
61.3
–
–
BeatFM-MusicFM
89.1
80.6
93.5
79.6
74.4
88.7
60.5
–
–
BeatMamba [ 28 ]
88.9
82.3
94.3
–
–
–
58.7
46.8
62.4
Table 3: Comparison with recent beat-tracking systems. BeatThis (direct) and CASM results were obtained from our experiments, while all other results are reported from their original papers.
Figure 2: Sensitivity to calibration data for CASM and DBN. Each point is one decoder configuration, calibrated on one of the 7/21/35 possible 1-/2-/4-fold subsets of SMC folds 1–7 (7F: all seven folds, a single configuration) and evaluated on SMC fold 0 and GTZAN. Because the subsets overlap, the boxes describe spread only. Dashed lines show the direct peak-picking performance as the reference.
School of Future Technology South China University of Technology Guangzhou, China · School of Intelligence Science and Technology Nanjing University Suzhou, China