Organizations: School of Information Science and Technology, Dalian Maritime University, Dalian, China · School of Automation, Beijing Institute of Technology, Beijing, China · School of Computer Science and Engineering, Northeastern University, Shenyang, China · School of Computer Science and Technology, Dalian University of Technology, Dalian, China
Large language models (LLMs) have shown promising potential for biomedical named entity recognition (BioNER) through instruction following and in-context learning. However, existing LLM-based BioNER methods still face two key limitations. First, retrieved demonstrations and external biomedical knowledge provide limited support for dataset-specific annotation semantics, leaving entity boundaries, type scopes, and annotation conventions ambiguous. Second, free-form generation lacks sufficient structural control, often leading to invalid formats, hallucinated mentions, duplicated entities, and boundary errors. To address these limitations, we propose GAMA, a guideline-augmented multi-agent framework for schema-as-code BioNER. GAMA first induces candidate annotation rules from labeled training instances and verifies them against annotated data to construct reliable dataset-specific guideline memory. Guided by these verified rules, a planning component generates ranked span-type hypotheses with rationales, and a coding component converts them into schema-constrained entity objects. A verification module then checks span grounding, type validity, and structural compliance, and performs dual-loop refinement to correct invalid or low-confidence predictions. Experiments on five widely used BioNER datasets with multiple LLM backbones show that GAMA consistently outperforms strong LLM-based baselines. Ablation and parameter analyses further verify the effectiveness of the proposed components.
Figures & tables
Fig. 1: Overview of the proposed GAMA framework. The framework instantiates a guideline-augmented multi-agent paradigm for schema-as-code biomedical named entity recognition. It consists of five stages: rule induction from training annotations, guideline memory construction, guideline-guided mention planning, schema-as-code entity generation, and dual-loop verification refinement for producing final structured entities.
Fig. 2: Code-based template for BioNER entities. span denotes the extracted biomedical mention, and entity_type denotes its predicted category.
Dataset
Labels
Train
Test
Entities
NCBI
1
5432
942
6881
BC2GM
1
12632
5065
24583
BC4CHEMD
1
31770
4240
84310
GENIA
5
16692
1854
50509
AnatEM
1
5861
3830
13701
TABLE I: Statistics of the BioNER datasets used in our experiments.
LLM
Method
NCBI
BC2GM
BC4CHEMD
GENIA
AnatEM
Llama3-8B
CodeIE
68.24
59.94
68.25
57.25
70.29
GPT-NER
62.89
52.97
65.21
48.28
64.98
CMAS
61.21
53.64
65.54
52.72
64.75
EICL
73.56
63.41
76.29
60.31
75.26
GAMA
74.62
66.35
76.74
62.45
75.95
Llama3-70B
CodeIE
73.25
65.23
72.41
63.51
75.15
TABLE II: Performance comparison of different methods across five BioNER datasets under different LLMs. The best results for each LLM-dataset combination are highlighted in bold .
LLM
Method
NCBI
BC2GM
BC4CHEMD
GENIA
AnatEM
Qwen2.5-14B
CodeIE
68.47
60.13
68.89
57.54
70.71
GPT-NER
63.08
53.25
65.54
48.62
65.19
CMAS
61.76
53.94
65.93
53.09
65.06
EICL
73.96
63.88
76.73
60.62
75.65
GAMA
75.12
66.91
77.39
62.94
76.42
Qwen2.5-72B
CodeIE
73.06
65.48
72.13
63.26
74.98
TABLE III: Performance comparison of different methods under multiple LLM backbones (Qwen2.5-14B, Qwen2.5-72B, GPT-3.5-turbo, and GPT-4o). The best results for each LLM group are highlighted in bold .
Llama3-70B
Ablation Setting
NCBI
BC2GM
BC4CHEMD
GAMA (full model)
79.26
70.69
81.68
w/o Guideline Modeling
75.49
66.27
77.42
w/o Rule Verification
76.21
66.98
77.95
w/o Planning Rationales
74.17
66.37
76.43
w/o Schema-as-Code
77.63
68.98
79.92
TABLE IV: Ablation studies on two different LLMs, evaluated on three BioNER datasets.
Seltasquare Seoul, Korea · Biomedical Research Institute, Seoul National University Hospital · Seoul National University College of Medicine Seoul, Korea
School of Information Sciences, University of Illinois Urbana-Champaign, 501 E Daniel St., Champaign, 61820, IL, USA · Division of Computational Health Sciences, Department of Surgery, University of Minnesota, 516 Delaware St SE, Minneapolis, 55455, MN, USA