eess.ASOct 7, 2026

CASM: Context-Aware Semi-Markov Post-Processor for Beat Tracking

Authors: Zhanhong He, Hanyu Meng, Yaolong Ju

Organizations: The University of Western Australia, Perth, Australia · The University of New South Wales, Sydney, Australia · Great Bay University, Dongguan, China

Abstract

In music beat tracking, model-predicted activations must be decoded into a discrete, musically coherent event sequence. Direct peak picking closely follows local evidence but can retain spurious peaks or miss weak beats. Dynamic Bayesian networks (DBNs), a widely used structured post-processor, improve sequence consistency under predefined global tempo, meter, and transition constraints, but their behavior can depend strongly on these settings. We introduce CASM, a context-aware semi-Markov decoder that instead conditions its temporal constraint on local activation evidence. CASM also accounts for ambiguity among competing periodic interpretations, including half- and double-tempo alternatives. Deterministic safeguards prevent implausible outputs and preserve beat-downbeat consistency. Applied to fixed activations from three neural beat trackers (BeatThis, MSCNN, and TCN), CASM improves temporal continuity while preserving event-level F1 across the GTZAN and SMC datasets, without backbone retraining or dataset-specific retuning. Further analysis shows that CASM is less sensitive than the DBN baseline to the composition of the calibration data.

Figures & tables

Explore similar work

Aug 5, 2026cs.SD

Masked diffusion enables coherent beat tracking

Current neural networks for beat tracking generate invalid outputs, such as consecutive downbeats and erratic tempo changes, even when these are not present in the training data. Heavy post-processing techniques can alleviate these problems, but the original cause of this inconsistent behaviour remains unknown. We hypothesise that it stems from inadequate modelling of multiple plausible output beat grids, resulting in an invalid mixture of competing interpretations. We propose a masked diffusion approach that properly models multiple outputs and enables the model to build coherent predictions through iterative inference. We devise three modifications to standard masked diffusion that enable its application to beat tracking: independent masking of beats and downbeats during training and inference, a balanced masking scheduler for inference, and peak-picking across inference steps. Our approach reduces erratic behaviours and improves beat-tracking performance.
May 12, 2026eess.AS

The SMC Blind Spot: A Failure Mode Analysis of State-of-the-Art Beat Tracking

Over the past two decades, the task of musical beat tracking has transitioned from heuristic onset detection algorithms to highly capable deep neural networks (DNN). Although DNN-based beat tracking models achieve near-perfect performance on mainstream, percussive datasets, the SMC dataset has stubbornly yielded low F-measure scores. By testing how well state-of-the-art models detect beats on individual tracks in the SMC dataset, we identify three distinct failure modes: octave errors, continuity errors, and complete tracking failure where all metrics fall below 0.3. We reveal that state-of-the-art models tend to generate "confident-but-wrong" activations. Furthermore, we show that the standard DBN's default minimum tempo of 55 BPM prevents it from inferring the correct tempo for 21% of SMC tracks, forcing double-tempo predictions on slow music. By exposing such fundamental oversights, we provide concrete directions for improving beat and downbeat detection, specifically emphasizing training data diversification and multi-hypothesis tempo estimation.
Jul 13, 2026cs.SD

BeatEdit: Symbolic Music Generation as Explicit Editing

Music creation is fundamentally a process of revision. Yet symbolic music generation remains dominated by paradigms that produce complete sequences from scratch, with limited support for selective modification. Edit-based methods have proven effective for text transformation tasks, but remain largely unexplored for symbolic music. We trace this absence to the representational level: conventional event-based music encodings lack the structural properties required by explicit music editing. In contrast, the BEAT encoding, a beat-grid-anchored representation originally designed for autoregressive generation, possesses structural properties amenable to editing. We propose BeatEdit, the first framework for symbolic music generation based on explicit edit operations, recasting generation as producing new content by editing a draft rather than synthesizing from scratch. BeatEdit comprises three complementary mechanisms along an axis of increasing edit density: per-token sequence tagging for error correction, iterative refinement for accompaniment editing, and tag-then-fill for segment completion. All these mechanisms share a single encoding and pre-trained backbone, achieving higher precision and perceptual quality than autoregressive and diffusion methods across all three tasks, while remaining efficient, with single-pass inference completing in under 100 ms. Cross-encoding evaluation further reveals that encoding design substantially influences editing effectiveness, with notable encoding-method interaction effects. Code is available at https://github.com/Haoyu-Gu/BeatEdit-code