EAGER: Enhancing Generative Event Extraction via Reinforcement Learning with Verifiable Rewards
Organizations: German Research Center for Artificial Intelligence (DFKI), Germany · Carl von Ossietzky Universität Oldenburg, Germany
Abstract
End-to-end event extraction remains challenging for large language models as it requires simultaneous identification of event triggers, classification of event types, and extraction of schema-grounded argument spans. We present EAGER, a reinforcement learning framework for generative event extraction that combines fine-grained verifiable rewards with Schema-Contrastive Advantage Estimation to alleviate advantage collapse under sparse binary rewards. Our reward design explicitly targets structural validity, extraction accuracy, groundedness, coverage, over-generation, and span precision. Experiments across seven benchmark datasets show that EAGER consistently outperforms prompting, supervised fine-tuning, and prior reinforcement learning baselines, achieving a substantial improvement over the strongest prior method. Results demonstrate that task-aligned verifiable rewards and contrastive advantage estimation substantially improve structured extraction.
Figures & tables
| Reward Component | Notation | Formulation | Role in Training |
|---|---|---|---|
| Validity | Binary: 1 if output parses into the required structured representation, 0 otherwise | Filters malformed outputs before any extraction scoring | |
| Extraction Accuracy | Macro-average F1 over trigger identification (TI), trigger classification (TC), argument identification (AI), argument classification (AC), AI + , and AC + | Captures canonical end-to-end extraction quality | |
| Groundedness | Verifies that predicted triggers and arguments appear verbatim in the source input; averages span-level support scores | Suppresses hallucinated spans and improves factual consistency | |
| Over-generation | Penalizes predictions that exceed the gold event and argument counts | Reduces false positives and limits excessive output length | |
| Coverage | Proportional recall over gold events and arguments, normalized to prevent inflation from over-prediction | Encourages completeness while balancing over-generation penalties | |
| Span Precision | Token-level Jaccard similarity with an additional penalty for oversized predicted spans | Rewards near-correct boundary predictions and discourages span inflation |
| Model | WikiEvents | PHEE | CASIE | GENIA11 | GENIA13 | MLEE | M2E2 | Avg. |
|---|---|---|---|---|---|---|---|---|
| Few-shot Prompting | ||||||||
| o1 Jaech et al. (2024) | 13.08 | 2.86 | 5.49 | 10.27 | 7.95 | 12.80 | 14.24 | 9.52 |
| GPT-4o Hurst et al. (2024) | 10.31 | 2.00 | 3.47 | 12.43 | 10.87 | 13.35 | 11.97 | 9.20 |
| GPT-5.4-mini Singh et al. (2025) | 10.58 | 18.86 | 5.38 | 11.95 | 10.56 | 10.70 | 13.86 | 11.70 |
| GPT-5.4 Singh et al. (2025) | 8.30 | 1.42 | 6.22 | 11.94 | 8.74 | 11.90 | 17.74 | 9.47 |
| GoLLIE-7B Sainz et al. (2024) | 13.65 | 30.44 | 7.47 | 12.58 | 9.69 | 8.49 | 32.14 | 16.35 |
| Setting/Rewards | Wiki | PHEE | CASIE | G11 | G13 | MLEE | M2E2 | Avg |
|---|---|---|---|---|---|---|---|---|
| ALL Rewards | 21.31 | 60.36 | 14.66 | 26.32 | 23.39 | 23.42 | 46.09 | 30.79 |
| + | 20.12 | 53.66 | 15.70 | 26.11 | 24.10 | 21.62 | 30.22 | 27.36 |
| + | 22.40 | 57.74 | 14.94 | 26.83 | 22.48 | 20.02 | 36.72 | 28.73 |
| + | 23.04 | 58.27 | 12.96 | 26.54 | 22.14 | 19.33 | 39.20 | 28.78 |
| + | 21.38 | 58.01 | 16.12 | 26.34 | 24.22 | 21.53 | 31.52 | 28.45 |
| + | 21.46 | 58.97 | 15.91 | 26.63 | 24.27 | 21.40 | 34.26 | 28.99 |
| Setting | Wiki | PHEE | CASIE | G11 | G13 | MLEE | M2E2 | Avg |
|---|---|---|---|---|---|---|---|---|
| GoLLIE-7B | 13.65 | 30.44 | 7.47 | 12.58 | 9.69 | 8.49 | 32.14 | 16.35 |
| w/o A.G. | 9.14 | 7.67 | 3.13 | 1.64 | 1.17 | 2.41 | 15.54 | 5.81 |
| EAGER | 21.31 | 60.36 | 14.66 | 26.32 | 23.39 | 23.42 | 46.09 | 30.79 |
| w/o A.G. | 19.43 | 43.66 | 8.34 | 17.32 | 16.96 | 12.25 | 40.33 | 22.61 |
| Dataset | Method | TI | TC | AI | AC | AI+ | AC+ |
|---|---|---|---|---|---|---|---|
| WikiEvents | All | 38.92 | 30.83 | 17.28 | 14.51 | 14.32 | 12.01 |
| 39.53 | 32.37 | 15.33 | 12.28 | 11.73 | 9.47 | ||
| + | 39.88 | 35.31 | 16.96 | 14.10 | 12.17 | 10.37 | |
| + | 39.22 | 35.37 | 18.22 | 16.12 | 15.25 | 14.05 | |
| + | 38.74 | 34.12 | 17.92 | 15.89 | 14.66 | 13.09 | |
| + | 41.53 | 34.65 | 15.83 | 13.52 | 12.20 | 10.52 |
| Setting | Wiki | PHEE | CASIE | G11 | G13 | MLEE | M2E2 | Avg |
|---|---|---|---|---|---|---|---|---|
| EAGER | 21.31 | 60.36 | 14.66 | 26.32 | 23.39 | 23.42 | 46.09 | 30.79 |
| w/o SCAE | 13.98 | 50.69 | 8.67 | 15.72 | 15.24 | 12.22 | 41.24 | 22.54 |
| EAGER + | 20.12 | 53.66 | 15.70 | 26.11 | 24.10 | 21.62 | 30.22 | 27.36 |
| w/o SCAE | 15.49 | 52.32 | 9.03 | 16.44 | 16.52 | 11.67 | 43.46 | 23.56 |
| Reward Added | Avg. Trig. | Avg. Arg. |
|---|---|---|
| All |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | WikiEvents | PHEE | CASIE | GENIA11 | GENIA13 | MLEE | M2E2 | Avg. |
|---|---|---|---|---|---|---|---|---|
| GoLLIE-7B w/o AG | 9.14 | 7.67 | 3.13 | 1.64 | 1.17 | 2.41 | 15.54 | 5.81 |
| GoLLIE-13B w/o AG | 10.07 | 7.78 | 4.00 | 2.16 | 0.99 | 2.88 | 22.31 | 7.17 |
| GoLLIE-34B w/o AG | 10.46 | 9.71 | 4.25 | 5.26 | 5.00 | 4.55 | 29.37 | 9.80 |
| Method | Model | GPUs | Batch Size | Epochs | Wall Time (h) |
|---|---|---|---|---|---|
| EAGER | GoLLIE-7B | 4 H100 | 4 | 1 | 4 |
| Dataset | #Docs | #Instances | #Event Types | #Events | #Arguments | Domain |
|---|---|---|---|---|---|---|
| WikiEvents | 245 | 565 | 50 | 598 | 5,501 | Wikipedia |
| CASIE | 999 | 1,375 | 5 | 8,469 | 22,575 | Cybersecurity |
| PHEE | 4,827 | 4,827 | 2 | 5,019 | 25,760 | Pharmacovigilance |
| GENIA2011 | 960 | 960 | 9 | 13,537 | 11,865 | Biomedical |
| GENIA2013 | 20 | 664 | 13 | 6,001 | 5,660 | Biomedical |
| MLEE | 262 | 286 | 29 | 6,575 | 5,958 | Biomedical |
| Acronym | Error Category | Definition |
| Event-level Errors | ||
| ME | Missing Events | A gold event instance is absent from the model’s output; the trigger and all its associated arguments are undetected. |
| EE | Extra Events | The model predicts an event instance with no corresponding gold event; a spurious trigger is generated that is not grounded in the annotation. |
| Argument-level Errors | ||
| MA | Missing Arguments | A gold argument span is not predicted for an otherwise correctly identified event; the event is detected but one or more of its argument slots are left unfilled. |
| EA | Extra Arguments | The model predicts one or more argument spans that have no corresponding gold argument; spurious fillers are generated beyond the gold annotation. |
| Method | WikiEvents | PHEE | CASIE | ||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| TI | TC | AI | AC | AI+ | AC+ | Avg. | TI | TC | AI | AC | AI+ | AC+ | Avg. | TI | TC | AI | AC | AI+ | AC+ | Avg. | |
| GPT-4o Hurst et al. (2024) | |||||||||||||||||||||
| GPT-5.4-mini Singh et al. (2025) | |||||||||||||||||||||
| GPT-5.4 Singh et al. (2025) | |||||||||||||||||||||
| Gollie-7B Sainz et al. (2024) | |||||||||||||||||||||
| Gollie-13B Sainz et al. (2024) | |||||||||||||||||||||
| Method | Genia2011 | Genia2013 | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| TI | TC | AI | AC | AI+ | AC+ | Avg. | TI | TC | AI | AC | AI+ | AC+ | Avg. | |
| GPT-4o Hurst et al. (2024) | ||||||||||||||
| GPT-5.4-mini Singh et al. (2025) | ||||||||||||||
| GPT-5.4 Singh et al. (2025) | ||||||||||||||
| Gollie-7B Sainz et al. (2024) | ||||||||||||||
| Gollie-13B Sainz et al. (2024) | ||||||||||||||
| Method | MLEE | M2E2 | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| TI | TC | AI | AC | AI+ | AC+ | Avg. | TI | TC | AI | AC | AI+ | AC+ | Avg. | |
| GPT-4o Hurst et al. (2024) | ||||||||||||||
| GPT-5.4-mini Singh et al. (2025) | ||||||||||||||
| GPT-5.4 Singh et al. (2025) | ||||||||||||||
| Gollie-7B Sainz et al. (2024) | ||||||||||||||
| Gollie-13B Sainz et al. (2024) | ||||||||||||||