Organizations: Department of Informatics (DI), Pontifícia Universidade Católica do Rio de Janeiro (PUC-Rio), Rio de Janeiro, RJ 22451-900, Brazil · Sorbonne Université, CNRS, LIP6, F-75005 Paris, France · CBAE, Universidade Federal do Rio de Janeiro (UFRJ), Rio de Janeiro, RJ 22250-020, Brazil
Most music-generation systems are still framed and evaluated primarily as producers of complete outputs, whereas composition often proceeds through successive revisions to a shared musical artifact. This paper studies a different use of a general-purpose instruction-following large language model: not as a one-shot music generator, but as a reusable operator over an evolving symbolic score. We formulate incremental composition as a sequence of operation-aware state transitions over persistent ABC notation, with explicit requirements on what each operation may change and what it must preserve. The interaction includes two artifact-initialization variants and three editing operations -- chord addition, inpainting, and transposition. We instantiate the formulation by adapting Llama 3.1 8B Instruct with Low-Rank Adaptation (LoRA) on 496,038 operation-aware dialogue records derived from Irish traditional music. The comparison with the unadapted model is used to test the feasibility of learning this interaction contract, not to claim novelty for fine-tuning itself. Across 500 dialogues per model (1,750 attempted output states), checker admission rises from 29.37% to 99.37%, while compliance conditional on admission rises from 0.7205 to 0.9798. Strict eligibility for reference-relative musical-feature analysis increases from 14 to 1,548 outputs, and Longest Common Subsequence analysis does not show a systematic increase in high-overlap sequences relative to held-out baselines under the specified protocol. The results support the technical feasibility of persistent, operation-aware symbolic editing with a general-purpose instruction LLM. They do not establish superior musical quality or human-AI co-creativity, which remain questions for musician-centered evaluation.
Figures & tables
Figure 1 : Three qualitative projections of the focused design space used to position related systems: (a) interaction continuity versus transformation scope, (b) control modality versus operation semantics, and (c) relation to the existing artifact versus preservation contract. Marker positions are explanatory, not metric scores or rankings. The dashed Suno/Udio arrow is a contextual product-design trend, not an academic comparison.
Figure 2 : Operation-aware state transition over a persistent ABC artifact. The requested operation determines both what may change and what must be preserved; the revised artifact becomes the state for the next turn.
Figure 3 : A minimal lead-sheet example showing the same musical material represented as ABC text (top) and as conventional notation (bottom).
Figure 4 : Denominator-aware evaluation architecture. Syntax admission is shared, after which operation reliability, reference-relative musical features, and corpus-relative memorization follow separate protocols and populations.
Figure 5 : Schematic boundary effect for non-negative pairwise distances. Standard Gaussian KDE assigns mass below zero; the Schuster mirror-image estimator reflects that contribution back onto the valid support.
State
Base adm.
FT adm.
Base comp.
FT comp.
Add Chords
.276
.996
.3910
.9702
Gen. before harm.
.300
.996
.7148
.9514
Inpaint + chords
.348
.980
.7942
.9904
Transpose + chords
.332
.984
.8221
.9796
Melody generation
.276
1.000
.7346
.9832
Melody inpainting
.256
1.000
.6944
.9915
Table 1 : Admission and conditional compliance (250 attempts per output state). FT = fine-tuned.
Figure 6 : Per-state decomposition of Table 1 . Admission and compliance both improve across all seven evaluated output states. Melody transposition is the clearest case where the base model retains some conditional competence once a response reaches the checker, yet still suffers from poor end-to-end admission.
Figure 7 : Population retained at successive evaluation stages, expressed as a percentage of the 1,750 attempts per model. “Strict feature eligible” is the pre-equalization musical-feature population.
Turn
Request
Observable state effect
1
Create in A Dorian, 6/8
Establishes the melody artifact.
2
Add chords
Inserts "Am" and "G" ; melody unchanged.
3
Inpaint bars 3–5
Rewrites target region; other bars retained.
4
Transpose to C Dorian
Transposes melody and chord symbols.
Table 2 : A four-turn evaluated dialogue showing persistent state transitions.
Figure 8 : Illustrative rendering of the four-turn persistent-artifact workflow summarized in Table 2 . Each turn produces a revised ABC artifact that becomes the direct input to the next request. The example is schematic, but it matches the evaluated semantics: creation establishes the artifact, chord addition preserves melody, inpainting rewrites only the targeted region, and transposition updates both melody and chord symbols together.
Figure 9 : Prototype of the extended NONOTO score editor illustrating the natural-language/ABC interaction workflow considered in this work.
Symbolic music evaluation for large language models remains fragmented across representations, datasets, and metrics. We introduce LilyBench, a LilyPond-based benchmark that jointly evaluates symbolic music generation and music understanding on the same family of open-weight LLMs. The benchmark includes a 200-prompt generation suite and ten understanding tasks adapted from ABC-Eval, covering syntax, metadata prediction, structural sequencing, and music recognition. Generation quality is evaluated using compile rate, MusPy descriptor distributions via Jensen-Shannon similarity, and LilyBERT-based Fréchet Music Distance (FMD). Experiments on four open-weight models show that executable LilyPond generation is achievable in zero-shot settings, while structural understanding tasks remain challenging despite strong performance on composer and genre recognition. Our experiments also reveal systematic disagreements between descriptor-based and embedding-based metrics, suggesting that symbolic music evaluation benefits from metric triangulation rather than single-score ranking. We release the benchmark, prompt bank, and evaluation code to support future research in symbolic music generation and understanding at https://github.com/CSCPadova/lilybench
Matteo Spanio, Mohammad Torabi, Andrea Poltronieri +1
Centro di Sonologia Computazionale, University of Padova, Padova, Italy · Music Technology Group, Universitat Pompeu Fabra, Barcelona, Spain
Large language models can produce superficially legal twelve-tone scores that collapse into degenerate textures. We introduce a neuro-symbolic harness that wraps a language-model proposer in a generate-verify-repair-trace loop with symbolic verification. The complete pipeline improves event-local consistency without claiming whole-piece legality. Across 40 controlled tasks and four paired models, constraint-checked delivery rises from 13.3% under raw generation to 48.1% with the harness; it abstains on the remaining 51.9% of runs. The pass rate of a narrower collision and serialisation-consistency check rises from 33.5% to 58.3%, while degeneracy remains near 0.05, including under adversarial prompting. A blinded evaluation by five experts also shows a descriptive aggregate preference for harness candidates over raw generation in adherence, perceived legality, coherence, and overall quality.
Congren Dai, Danni Zhao, Enyang Liu +7
Central Conservatory of Music · Tsinghua University
Generative music systems can now produce impressive audio from text prompts, but audio outputs are difficult to inspect, edit, and diagnose as musical structure. We introduce Libretto, an agent-facing framework for symbolic music generation and revision. Libretto uses an LLM-native grammar with explicit onset slots, voices, and bar-level organization, then evaluates each piece in a corpus-calibrated statistical space over rhythm, harmony, melody, texture, form, and variation. The same structural axes support retrieval, diagnosis, copy-risk control, and iterative self-revision. Across gap filling, reference-guided full-piece generation, gradual morphing, and educational music generation, Libretto turns symbolic music from a raw token sequence into a measurable and editable object for language-model agents.