On the (In)effectiveness of AMR Augmentation for Large Language Models
Organizations: Saarland University · University of Trento
Abstract
While Abstract Meaning Representation (AMR) has historically improved performance on a range of NLP tasks, the benefit---or lack thereof---of AMR augmentation for modern LLMs is thus far unclear. In this paper, we attempt to reproduce recent work that reported substantial downstream gains from AMR augmentation, finding that these are likely due to specific choices in the experimental settings used: using a consistent and unified protocol for hyperparameter selection, we observe that text-only baselines consistently match or exceed the performance of AMR-augmented models. To investigate this null result, we introduce a perplexity-based probe measuring the degree to which AMR provides an LLM with supplemental relational knowledge not already available to the model. We find that AMR augmentation does not help LLMs improve their understanding of relational content in the sentence, indicating that augmenting these models with AMR offers no clear benefit on downstream tasks.
Figures & tables
| Format | Example |
|---|---|
| Text-only | { sys } You are an expert in machine translation… |
| { user } Obama receives Netanyahu | |
| { asst } Obama empfängt Netanyahu | |
| +AMR | { sys } You are an expert in machine translation… |
| { user } Sentence: Obama receives Netanyahu | |
| AMR: (r / receive-01 :ARG0 (p / person :name (n / name :op1 "Obama")) :ARG1 (p2 / person :name (n2 / name :op1 "Netanyahu"))) |
| Model | Training | Setup | WiC | SST | SNLI | PM | PAWS | AGN | SPD | CNL | WMT |
|---|---|---|---|---|---|---|---|---|---|---|---|
| (F1) | (F1) | (F1) | (F1) | (F1) | (F1) | (EM) | (F1) | (BLEU) | |||
| Qwen3-8B | Joint | Text | 0.3 | 1.2 | 0.8 | 2.1 | 0.7 | 0.5 | 1.1 | 0.3 | 0.5 |
| +AMR | 0.4 | 1.9 | 1.2 | 2.5 | 0.4 | 0.3 | 1.2 | 0.6 | 0.8 | ||
| +AMR +Inter. | 1.7 | 0.3 | 1.3 | 4.8 | 0.7 | 0.3 | 1.5 | 1.6 | 0.7 | ||
| Individual | Text | 2.1 | 0.5 | 1.2 | 2.5 | 0.3 | 0.8 | 1.2 | 0.5 | 0.2 | |
| +AMR | 1.7 | 1.4 | 1.3 | 1.4 | 0.4 | 0.5 | 1.3 | 0.3 | 0.1 |
| Dataset | Llama | Qwen | Pooled |
|---|---|---|---|
| Individual training | |||
| WMT | 0.370 | 0.067 | 0.985 |
| PAWS | 0.166 | 0.273 | 0.064 |
| WiC | 0.954 | 0.631 | 0.717 |
| SPIDER | 0.790 | 0.614 | 0.901 |
| AGNews | 0.590 | 0.297 | 0.211 |
| Dataset | Task | Train | Dev | Test |
|---|---|---|---|---|
| RAMS | Event Argument Extraction | 7000 | 849 | 817 |
| CNN/DM | Abstractive Summarization | 5000 | 500 | 500 |
| ANLI | Natural Language Inference | 7000 | 700 | 1200 |
| LogiQA | Reading Comprehension | 3828 | 478 | 480 |
| Model | Training | Setup | RAMS | LogiQA | ANLI | CNN |
|---|---|---|---|---|---|---|
| (F1) | (Acc.) | (F1) | (ROUGE-L) | |||
| Llama-3.1-8B | Joint | Text | 50.0 ±0.6 | 54.9 ±0.5 | 88.5 ±0.5 | 31.5 ±0.4 |
| +AMR | 49.1 ±0.5 | 53.6 ±0.5 | 88.1 ±0.4 | 31.2 ±0.4 | ||
| +AMR +Inter. | 49.5 ±0.7 | 53.8 ±0.5 | 86.8 ±0.6 | 31.4 ±0.4 | ||
| Individual | Text | 49.4 ±0.7 | 55.7 ±1.5 | 87.7 ±0.9 | 31.0 ±0.7 | |
| +AMR | 50.1 ±0.6 | 57.2 ±0.7 | 87.4 ±0.5 | 30.7 ±0.5 |
| Input Type | User Prompt |
|---|---|
| Text-only | Input Sentence: ‘‘The NBA season of 1975--76 was the 30th season of the National Basketball Association.’’ |
| Text+AMR | Input Sentence: ‘‘The NBA season of 1975--76 was the 30th season of the National Basketball Association.’’ Input AMR: (s / season :ord (o / ordinal-entity :value 30) :poss (l / league :name (n / name :op1 ...))) |
| Text+AMR-nodes | Input Sentence: ‘‘The NBA season of 1975--76 was the 30th season of the National Basketball Association.’’ Supplement: season, ordinal-entity, 30, league, name, National, Basketball, Association, date-interval, date-entity, 1975, date-entity, 76 |
| Model | Text | +AMR | +AMR-Nodes |
|---|---|---|---|
| Llama | |||
| Base | |||
| Text-FT | |||
| AMR-FT | |||
| Inter-FT | |||
| Qwen | |||
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Task | Zhang et al. 2025 | Our Experiment | ||||
|---|---|---|---|---|---|---|---|
| Train | Val | Test | Train | Val | Test | ||
| PAWS | Paraphrase Detection | 10,000 | N/A | 8,000 | 10,000 | 1,000 | 8,000 |
| SNLI | Textual Entailment | 10,000 | N/A | 10,000 | 10,000 | 1,000 | 10,000 |
| WMT16 | Translation | 10,000 | N/A | 5,999 | 10,000 | 1,000 | 5,999 |
| CoNLL2003 | Named Entity Recog. | 10,000 | N/A | 3,453 | 10,000 | 1,000 | 3,453 |
| LOGIC | Logical Fallacy Det. | 10,000 | N/A | 2,449 | N/A | N/A | N/A |
| Train | Dataset | Setup | LR |
| Indiv. | PubMed | Text | |
| AMR | |||
| AMR-i. | |||
| WMT | Text | ||
| AMR | |||
| AMR-i. |
| Train | Dataset | Setup | LR |
| Indiv. | PubMed | Text | |
| AMR | |||
| AMR-i. | |||
| WMT | Text | ||
| AMR | |||
| AMR-i. |
| Parameter | Value |
|---|---|
| LLM Related | |
| Attention impl. | flash_attention_2 |
| Dtype | bfloat16 |
| LoRA Configuration | |
| Rank ( ) | 64 |
| Alpha | 128 |
| Training | Setup | WiC | SST | SNLI | PM | PAWS | AGN | SPD | CNL | WMT |
|---|---|---|---|---|---|---|---|---|---|---|
| (F1) | (F1) | (F1) | (F1) | (F1) | (F1) | (EM) | (F1) | (BLEU) | ||
| Joint | Text | 76.2 | 96.1 | 92.1 | 67.4 | 93.9 | 92.6 | 59.1 | 93.0 | 30.2 |
| +AMR | 78.1 | 96.7 | 92.0 | 73.9 | 93.5 | 92.1 | 61.5 | 92.4 | 31.3 | |
| Individual | Text | 78.0 | 96.7 | 92.2 | 83.9 | 94.2 | 93.6 | 59.4 | 92.9 | 31.3 |
| +AMR | 78.5 | 97.0 | 92.0 | 85.2 | 93.7 | 93.2 | 58.8 | 93.6 | 28.2 |
| { System } Generate a sentence for the given Abstract Meaning Representation. |
| { User } (r / receive-01 :ARG0 (p / person :name :op1 ‘‘Obama’’) :ARG1 (p2 / person :name ‘‘Netanyahu’’)) |
| { Assistant } Obama receives Netanyahu |
| Model Strategy | BLEU Score |
|---|---|
| LLaMA3.1-8B-Ins | |
| Phase 1 checkpoint | 46.0 |
| Phase 2 checkpoint | 43.9 |
| Qwen3-8B | |
| Phase 1 checkpoint | 45.7 |
| Phase 2 checkpoint | 43.1 |
| Train | Model | Dataset | Setup | LR |
| Indiv. | Llama | RAMS | Text | |
| AMR | ||||
| AMR-i. | ||||
| LogiQA | Text | |||
| AMR | ||||
| AMR-i. |
| Parameter | Value |
|---|---|
| LLM Related | |
| Attention impl. | flash_attention_2 |
| Dtype | bfloat16 |
| LoRA Configuration | |
| Rank ( ) | 64 |
| Alpha | 128 |
| Strategy | Training | RAMS | LogiQA | ANLI | CNN |
|---|---|---|---|---|---|
| (F1) | (Acc.) | (F1) | (R-L) | ||
| GNN | Joint | 50.2 | 54.6 | 88.8 | 31.0 |
| Indiv. | 50.7 | 57.5 | 88.1 | 31.2 | |
| AMRBART | Joint | 49.5 | 56.9 | 88.7 | 32.1 |
| Indiv. | 49.4 | 56.9 | 87.7 | 31.2 |
| Strategy | Training | WiC | SST | SNLI | PM | PAWS | AGN | SPD | CNL | WMT |
|---|---|---|---|---|---|---|---|---|---|---|
| (F1) | (F1) | (F1) | (F1) | (F1) | (F1) | (EM) | (F1) | (BLEU) | ||
| GNN | Joint | 74.0 | 95.8 | 90.9 | 71.3 | 93.2 | 88.8 | 53.4 | 90.9 | 27.7 |
| Individual | 74.7 | 96.2 | 91.0 | 74.4 | 93.5 | 93.0 | 52.1 | 92.4 | 26.6 | |
| AMRBART | Joint | 72.3 | 96.0 | 91.0 | 81.7 | 93.7 | 74.9 | 52.8 | 93.0 | 26.1 |
| Individual | 75.0 | 96.3 | 91.5 | 84.4 | 93.9 | 91.4 | 49.8 | 92.2 | 28.4 |
| Type of Prompt | System Prompt | User Prompt |
|---|---|---|
| Text-only | Task: Deconstruct the following sentence into its underlying logical structure. Describe the meaning by identifying core events, the entities involved, and how they relate to one another. Example: Input Sentence: "The NBA season of 1975 -- 76 was the 30th season of the National Basketball Association ." Output: This refers to a season of the National Basketball Association. It was the 30th season of the league. The season took place during the 1975-76 time period. | Sentence: ‘‘Captain’’ was hulked in 1739 , and eventually broken up in 1762 . |
| Text + AMR | Task: Deconstruct the following sentence into its underlying logical structure. Describe the meaning by identifying core events, the entities involved, and how they relate to one another. You may use the provided AMR (Abstract Meaning Representation) as a supplement to guide you through generation. Example: Input Sentence: "The NBA season of 1975 -- 76 was the 30th season of the National Basketball Association ." Input AMR: (s / season | Sentence: ‘‘Captain’’ was hulked in 1739 , and eventually broken up in 1762 . AMR: (a / and xxxxx :op1 (h / hulk-01 xxxxxxxxxx :ARG1 (s / ship xxxxxxxxxxxxxxx :name (n / name xxxxxxxxxxxxxxxxxx :op1 "Captain")) xxxxxxxxxx :time (d / date-entity xxxxxxxxxxxxxxx :year 1739)) xxxxx :op2 (b / break-up-08 xxxxxxxxxx :ARG1 s xxxxxxxxxx :time (e / eventual) xxxxxxxxxx :time (d2 / date-entity xxxxxxxxxxxxxxx :year 1762))) |
| Text + AMR nodes | Task: Deconstruct the following sentence into its underlying logical structure. Describe the meaning by identifying core events, the entities involved, and how they relate to one another. You may use the supplement provided if you find it helpful. Example: Input Sentence: "The NBA season of 1975 -- 76 was the 30th season of the National Basketball Association ." Supplement: season, ordinal-entity, 30, league, name, "National", "Basketball", "Association", date-interval, date-entity, 1975, date-entity, 76 Output: This refers to a season of the National Basketball Association. It was the 30th season of the league. The season took place during the 1975-76 time period. | Sentence: ‘‘Captain’’ was hulked in 1739 , and eventually broken up in 1762 . Supplement: and, hulk-01, ship, name, "Captain", date-entity, 1739, break-up-08, s, eventual, date-entity, 1762 |