ARIA: Audio-Driven Melody-Tone Relation Modeling for Cantonese Lyric Authoring
Organizations: The Hong Kong University of Science and Technology (Guangzhou)
Abstract
Cantonese lyric writing requires close alignment between lexical tones and melodic pitch. Existing melody-guided lyric generation methods typically rely on symbolic melody to generate lyrics. However, in real songwriting scenarios, melodies are often expressed as raw singing audio or hummed recordings, where pitch is implicit, noisy, and unstructured, making these methods difficult to apply directly. To address this limitation, we propose ARIA, a two-stage audio-driven melody-tone relation modeling framework for Cantonese lyric authoring that generates Cantonese lyrics from singing recordings with provided character-level timestamps. Specifically, we first design a Tri-Stream Relation-Aware Tone Estimator (TRATE) to predict 0243 sequences from timestamped singing audio by modeling multi-stream acoustic cues and relational tonal structure. We then propose a Decoupled Retrieval-Augmented Tone-Conditioned Lyric Generator (DRA-TCLG) to generate fluent lyrics conditioned on predicted tonal plans with retrieval-enhanced lexical guidance. Moreover, we construct a large-scale aligned audio-Jyutping-0243 dataset from real Cantonese singing recordings to support this new task. Experimental results demonstrate that ARIA achieves strong performance in both 0243 prediction and tone-consistent lyric generation, validating the effectiveness of the proposed framework.
Figures & tables
| Model | Acc | F1 | NED | Params | |
|---|---|---|---|---|---|
| Rule RQ | 0.373 | 0.360 | 0.749 | 0.583 | 0 |
| Rule AQ | 0.370 | 0.357 | 0.747 | 0.585 | 0 |
| Qwen Rel. | 0.501 | 0.471 | 0.835 | 0.460 | 4.57B / 32.46M tr. |
| Qwen 3-Br. | 0.565 | 0.537 | 0.844 | 0.398 | 4.57B / 32.46M tr. |
| TRATE | 0.644 | 0.615 | 0.885 | 0.347 | 2.37M |
| Method | Alignment | Diversity | ||||||
|---|---|---|---|---|---|---|---|---|
| Harmony | Consistency | Avg Sim | Min Sim | MaD1 | MaD2 | MiD1 | MiD2 | |
| SongNet | 0.4846 | 0.0375 | 0.4911 | 0.0479 | 0.9354 | 0.9862 | 0.0253 | 0.3341 |
| SmBART | 0.8186 | 0.6755 | 0.4277 | 0.0008 | 0.8830 | 0.9410 | 0.0297 | 0.4258 |
| ToneCraft | 0.9684 | 0.9347 | 0.5036 | 0.0541 | 0.9552 | 0.9855 | 0.0134 | 0.1403 |
| DRA-TCLG (Ours) | 0.9740 | 0.9421 | 0.5053 | 0.0193 | 0.9544 | 0.9944 | 0.0190 | 0.2967 |
| Method | Alignment | Diversity | ||||||
|---|---|---|---|---|---|---|---|---|
| Harmony | Consistency | Avg Sim | Min Sim | MaD1 | MaD2 | MiD1 | MiD2 | |
| 0243 -to-Lyrics (reference 0243 input) | ||||||||
| SongNet | 0.5000 | 0.0353 | 0.4699 | 0.1034 | 0.9487 | 0.9894 | 0.1896 | 0.7352 |
| SmBART | 0.8277 | 0.6373 | 0.4138 | 0.0877 | 0.9037 | 0.9524 | 0.2273 | 0.7849 |
| ToneCraft | 0.9747 | 0.9466 | 0.4957 | 0.0997 | 0.9612 | 0.9882 | 0.0945 | 0.3763 |
| DRA-TCLG (Ours) | 0.9841 | 0.9687 | 0.4849 | 0.1100 | 0.9630 | 0.9951 | 0.1345 | 0.6246 |
| Metric | Method | ||
|---|---|---|---|
| Base | One-Stage | Two-Stage | |
| Alignment | |||
| Harmony | 0.9104 | 0.6449 | 0.9740 |
| Consistency | 0.7619 | 0.1579 | 0.9421 |
| Diversity | |||
| Avg Sim | 0.5269 | 0.4912 | 0.5053 |
| Metric | DRA-TCLG | ToneCraft | SmBART | SongNet |
|---|---|---|---|---|
| Tone–Mel. | 3.185 | 2.967 | 2.654 | 2.029 |
| Rhythm | 3.256 | 2.998 | 2.671 | 2.152 |
| Canto. | 3.233 | 3.167 | 2.710 | 2.077 |
| Lyric | 3.279 | 3.054 | 2.756 | 2.202 |
| Sing. | 3.167 | 3.015 | 2.627 | 2.027 |
| Pref. | 160 | 142 | 89 | 22 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Tone | Category | Chao | Singing tier | 0243 |
|---|---|---|---|---|
| 1 | Yin level | 55/53 | high | 3 |
| 2 | Yin rising | 35/25 | high | 3 |
| 3 | Yin departing | 33 | mid-high | 4 |
| 4 | Yang level | 11/21 | low | 0 |
| 5 | Yang rising | 13/23 | mid-high | 4 |
| 6 | Yang departing | 22 | mid | 2 |
| Tone | Strict | 0243 | Rationale |
|---|---|---|---|
| 1 | 3 | 3 | Same high tier as tone 2 |
| 2 | 9 | 3 | Same high tier as tone 1 |
| 3 | 4 | 4 | Same mid-high tier as tone 5 |
| 4 | 0 | 0 | Unique low tier |
| 5 | 5 | 4 | Same mid-high tier as tone 3 |
| 6 | 2 | 2 | Unique mid tier |
| Stage | Main risk | What we do |
|---|---|---|
| Initial alignment | Pickup notes, melisma, and ornamentation blur where a lyric character should begin and end in the audio. | We inspect sampled examples and use them to monitor alignment quality instead of assuming the raw alignment output is always reliable. |
| Abnormal-data filtering | Some examples have durations that are implausibly long, implausibly short, or clearly mismatched with the lyric character count. | We remove these abnormal items before forming the final dataset so that TRATE is not trained on obvious alignment failures. |
| Jyutping conversion | The final label sequence depends on having a stable tone category for each aligned lyric character. | Once the lyrics are aligned, we obtain Cantonese Jyutping on the text side and use it as the basis for deterministic label derivation. |
| 0243 derivation | Directly ranking pitch from audio would introduce extra ambiguity and subjective judgment. | We map Jyutping tone categories to 0243 deterministically, so no further manual verification of the final 0243 sequence is required. |
| Field | Description |
|---|---|
| sample_id | Line-level sample identifier |
| song_id | Source song identifier |
| line_id | Line index within the song |
| text_length | Number of lyric characters |
| char_bounds | Character start time and duration |
| label_0243 | Token-level 0243 labels |
| Criterion | Why 0243 helps | Implication for ARIA |
|---|---|---|
| Important | It preserves the most rigid Cantonese lyric constraint: compatibility between lexical tone tier and melodic motion. | The downstream generator does not lose the core “fit to melody” condition at the representation boundary. |
| Effective | It matches long-standing lyric-writing practice and recent harmony-aware Cantonese generation models. | The control signal is interpretable, linguistically motivated, and aligned with real authoring behavior. |
| Efficient | Four discrete classes are lower-entropy and more transposition-robust than raw audio, continuous F0, or full note strings. | Annotation is easier to standardize, search space is smaller, and conditioning an LLM becomes simpler and more stable. |
| Real-world lyric-writing step | Prior knowledge used by human lyricists | Corresponding ARIA design | Effect on the model |
|---|---|---|---|
| Listen to a demo and locate singable slots | Musical timing, phrase boundaries, and character-level rhythm. | TRATE takes singing audio with character boundaries . | The model predicts one controllable label per lyric position rather than a frame-level F0 trace. |
| Estimate tone-compatible melody tiers | Cantonese tones are organized by relative pitch height and contour, not only by absolute frequency. | TRATE uses absolute, contour, and line-internal relative feature streams. | Acoustic, melodic, and sentence-level evidence are separated before fusion. |
| Judge neighboring tone–melody fit | Lyricists compare adjacent syllables and avoid implausible large jumps when possible. | TRATE uses relation-aware attention and pairwise ordering heads to model short-range melodic relations. | The target is treated as relational, matching the ordinal nature of 0243 . |
| Use 0243 as a draftable authoring plan | The expert 0243 method collapses Cantonese tones into four practical singing tiers. | TRATE predicts level-4 labels and maps them deterministically to 0243 ; structured heads regularize register and contour factors. | Cantonese tonal knowledge enters supervision without giving the model lyric text at inference time. |
| Search and revise lexical candidates | Lyricists consult tone-compatible dictionaries and allow near alternatives for content. | DRA-TCLG conditions on , retrieves from a 0243 –lyrics dictionary, and refines candidates with an LLM. | The generator receives an explicit, human-readable lexical-planning signal instead of an opaque audio embedding. |
| Noise | Acc | F1 | NED | |
|---|---|---|---|---|
| Clean | 0.6459 | 0.6173 | 0.8840 | 0.3462 |
| ms | 0.6447 | 0.6167 | 0.8853 | 0.3465 |
| ms | 0.6391 | 0.6102 | 0.8837 | 0.3515 |
| ms | 0.6178 | 0.5871 | 0.8773 | 0.3704 |
| Feature | Definition | Role |
|---|---|---|
| mean_log_f0 | Mean log-F0 over frames assigned to the character. | Direct token-window pitch height before comparison with other tokens. |
| median_log_f0 | Median log-F0 over the character window. | Robust absolute height summary for the token itself. |
| std_log_f0 | Standard deviation of log-F0 in the character window. | Local acoustic dispersion, not normalized by line context. |
| min_log_f0 | Minimum log-F0 in the character window. | Lower bound of the token’s observed pitch region. |
| max_log_f0 | Maximum log-F0 in the character window. | Upper bound of the token’s observed pitch region. |
| q25_log_f0 | 25th percentile of character-window log-F0. | Absolute lower-quartile pitch statistic within the token. |
| Feature | Definition | Role |
|---|---|---|
| contour_log_f0 [0:15] | Sixteen evenly resampled log-F0 points across the character window. | Preserves pointwise within-character shape rather than collapsing the token into scalar statistics. |
| contour_voiced_mask [0:15] | Sixteen pointwise voiced/unvoiced indicators aligned to the resampled contour. | Attaches reliability evidence to each contour point. |
| contour_filled_mask [0:15] | Sixteen pointwise indicators showing where contour values were imputed. | Marks missing or filled trajectory regions in the contour sequence. |
| token_valid_voiced | Indicator that the character window contains native voiced frames before filling. | Provides a summary reliability flag for the extracted contour trajectory. |
| Feature | Definition | Role |
|---|---|---|
| robust_z_median | Character median log-F0 minus line median, divided by the line interquartile range. | Defines height relative to the current lyric line. |
| zscore_median | Character median log-F0 minus line mean, divided by line standard deviation. | Provides line-normalized height; this is the scalar used by Rule RQ. |
| percent_rank_median | Stable percent rank of character median log-F0 among tokens in the line. | Captures ordinal line-internal height independent of absolute register. |
| prev_interval | Current median log-F0 minus previous-token median log-F0. | Encodes adjacent-token pitch relation to the left context. |
| next_interval | Next-token median log-F0 minus current median log-F0. | Encodes adjacent-token pitch relation to the right context. |
| prev_higher_flag | Indicator that the current token is higher than the previous token. | Represents directional relation between neighboring tokens. |
| Model | Acc | F1 | NED | |
|---|---|---|---|---|
| Rule RQ | 0.399 | 0.436 | 0.593 | 0.625 |
| Rule AQ | 0.394 | 0.430 | 0.608 | 0.626 |
| Qwen Rel. | 0.460 | 0.489 | 0.671 | 0.544 |
| Qwen 3-Br. | 0.394 | 0.409 | 0.658 | 0.614 |
| TRATE | 0.454 | 0.471 | 0.677 | 0.550 |
| Case | Input plan | GT hit | Output use | Plan match |
|---|---|---|---|---|
| Use | 2433334 | Yes | 2 | Exact |
| Non-use | 22342433 | Yes | 0 | Exact |