Overview of BioASQ 2026: The fourteenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering
Authors: Anastasios Nentidis, Georgios Katsimpras, Anastasia Krithara, Martin Krallinger, Miguel Rodríguez-Ortega, Eduard Rodriguez-López, Natalia Loukachevitch, Igor Rozhkov, +13 more
Organizations: National Center for Scientific Research “Demokritos”, Athens, Greece · Barcelona Supercomputing Center, Barcelona, Spain · Lomonosov Moscow State University, Russia · Artificial Intelligence Research Institute, Russia · HSE University, Russia · Aristotle University of Thessaloniki, Greece · Northwell Health, New Hyde Park, New York, USA · Archimedes, Athena Research Center, Greece · University of Padua, Italy
This paper presents an overview of the fourteenth edition of the BioASQ challenge, organized in the context of the Conference and Labs of the Evaluation Forum (CLEF) 2026. BioASQ is an international challenge series that supports progress in biomedical language processing tasks ranging from semantic indexing and information extraction to question answering and summarization. In 2026, BioASQ included six shared tasks: a) Task 14b on biomedical semantic question answering. b) Task Synergy14 on question answering for developing biomedical top- ics. c) Task MultiClinSum-2 on multilingual clinical summarization. d) Task BioNNE-R on extracting relations between nested named entities in Russian and English. e) Task ELCardioCC on clinical coding in cardiology. f) Task GutBrainIE on gut-brain interplay information extrac- tion. Across these six tasks, 87 distinct teams participated, submitting more than 1000 runs overall. As in previous editions, several submissions reached competitive performance, reflecting the continued progress of state-of-the-art methods across biomedical language processing tasks.
Figures & tables
Batch
Size
Yes/No
List
Factoid
Summary
Documents
Snippets
Train
5729
1541
1130
1695
1363
10.74
13.97
Test 1
80
17
21
23
19
5.25
12.41
Test 2
80
21
13
20
26
3.64
7.98
Test 3
60
11
17
17
15
3.40
6.95
Test 4
60
16
20
11
13
3.78
8.35
Total
6009
1606
1201
1766
1436
10.43
13.75
Table 1: Statistics on the training and test datasets of task 14b. The numbers for the documents and snippets refer to averages per question.
Mean Word Count
Mean Sent. Count
Language
Nr. Pairs
Full case
Summary
Full case
Summary
Catalan (ca)
27699
604.80
118.56
27.49
5.59
Czech (cs)
27488
469.89
86.37
30.78
5.73
Danish (da)
27652
485.14
93.76
28.75
5.70
German (de)
27424
510.20
96.89
30.47
5.82
Greek (el)
26749
547.05
100.84
29.53
5.62
Table 2: Descriptive statistics per language sub-track of the MultiClinSum-2 dataset. For each language, the number of case-summary pairs and mean word and sentence counts are reported for both full clinical cases and their corresponding reference summaries.
Split
Docs
Entities
Rels
[RU]
train
716
37,879
23,773
dev
50
3,210
2,521
test
154
9,339
7,078
[EN]
train
55
3,861
3,193
Table 3: BioNNE-R Shared Task general document statistics. Entities are annotated entity mentions; relations are gold relation annotations.
Collection
# Docs
# Entities
Ents/Doc
# Rels
Rels/Doc
Train Gold
639
20530
32.13
8556
13.39
Train Silver
1310
41409
31.61
21523
16.43
Train Bronze
2972
89987
30.28
29692
9.99
Development Set
80
2521
31.51
1261
15.76
Test Set
80
2850
35.62
1285
16.06
Table 5: Dataset statistics for GutBrainIE .
Team
EN
ES
FR
PT
IT
RU
CA
NO
DA
RO
DE
EL
NL
CS
SV
Total
ixa-sum
5
5
2
2
2
—
5
—
—
2
2
—
2
—
—
27
MediScribes
5
—
—
—
—
—
—
—
—
—
—
—
5
—
—
10
NLP4Health
3
3
3
2
—
—
—
—
—
—
—
—
—
—
—
11
DACHausa
1
1
1
1
1
1
1
—
1
1
1
1
1
1
1
14
InfoLab-FEUP
3
—
—
3
—
—
—
—
—
—
—
—
—
—
—
6
PMH-Aqeel
1
1
—
1
1
1
1
—
—
—
—
—
—
—
1
7
Table 6: Participation overview of MultiClinSum-2 task per team and language subtrack. Numbers indicate runs submitted; — indicates no participation.
Resp. type (Phase)
Quest. type
Official measure
Documents (A)
All
Mean Average Precision (MAP)
Snippets (A)
All
F1 (based on character overlaps)
List
F1
Exact ans. (A+ & B)
Yesno
macro F1 on “yes” & “no” classes
Factoid
Mean Reciprocal Rank (MRR)
Ideal ans. (A+ & B)
All
Manual scores for precision, recall, repetition, readability
Table 7: The evaluation measures for task 14b per response type and question type [ 39 ] .
R
Qs
AR
Top MAP
Top F1 Snip.
Top MRR
Top F1 list
Top macro-F1
1
63
0
0.538
0.340
-
-
-
2
66
42
0.337
0.262
0.545
1.000
1.000
3
64
55
0.249
0.390
0.394
0.818
0.771
4
48
47
0.315
0.197
0.500
0.900
0.890
Table 10: The number of questions (Qs) and “Answer Ready” (AR) questions and the top system performance per round (R) in Task Synergy14. Retrieval of documents (Top MAP) and snippets (Top F1). Generation of exact factoid (Top MRR), list (Top F1), and yes/no (Top macro-F1) answers.
Lang.
Team
Run
R-1 (R-2)
R-L
BERT Score
Faith.
Com.
Fl.
Cons.
EN
ixa-sum
2
0.44 (0.22)
0.32
0.88
0.73
0.83
0.70
0.70
NLP4Health
1
0.38 (0.16)
0.27
0.87
0.76
0.94
0.70
0.70
ES
ixa-sum
2
0.47 (0.23)
0.31
0.88
0.71
0.80
0.70
0.65
NLP4Health
1
0.41 (0.18)
0.27
0.87
0.72
0.84
0.70
0.63
PT
NLP4Health
1
0.38 (0.16)
0.26
0.87
0.74
0.89
0.70
0.64
ixa-sum
1
0.38 (0.15)
0.26
0.87
0.74
0.91
0.70
0.65
Table 11: Top-2 participant teams per MultiClinSum-2 sub-track, ranked by BERTScore. Developed baseline system results are not included in the ranking (R-1 = ROUGE-1, R-2 = ROUGE-2, R-L = ROUGE-Lsum, Faith. = Faithfulness, Com. = Completeness, Fl. = Fluency, Cons. = Consistency).
Team
English
Russian
Bilingual
ELiRF-UPV
0.5060
0.4717
0.4540
LODAC-NII
0.5016
0.5283
0.5015
aakobiakova
0.4951
0.4403
0.4716
HSE NLP
0.4733
0.5057
0.4527
rabiaozdemir
0.4620
–
–
savvafq
0.4570
0.4829
0.4849
Table 12: Official BioNNE-R results: macro-F1 of each team’s best submission per subtask. Best result per subtask in bold; – indicates no submission.
Team
System
Recall
Precision
Micro-F1
stanimeros
ensemble_metaheuristic_p4_winner…
0.850979
0.882995
0.866691
Georgios_1
Baymax_submission_v1
0.865845
0.861783
0.863809
LSI_UNED
test_set_NER_greekuncased_greek…
0.818709
0.841595
0.829994
alikibliona
submission_A_reranker
0.818347
0.824626
0.821474
elcardiocc
baseline_bert_base_uncased_greek
0.659173
0.903579
0.762264
Pakakisd
max_micro_curriculum
0.809282
0.720000
0.762035
Table 13: Performance of participating systems in the ELCardioCC 2026
Team ID
Run name
Precision
Recall
F1
TWIX [ 43 ]
merged6
0.9548
0.8058
0.8740
NightSun [ 36 ]
run21
0.8756
0.8110
0.8420
Graphwise [ 59 ]
18
0.8613
0.8169
0.8385
TEXA [ 69 ]
1
0.7980
0.8302
0.8138
unibuc-bionlp-bp
1
0.7816
0.8425
0.8109
SMTE [ 45 ]
R1
0.8109
0.7932
0.8019
Table 14: GutBrainIE performance metrics of each team’s top run beating the baseline for NER. The best result is in bold, the second-best is underlined (micro-averaged).
Team ID
Run name
Precision
Recall
F1
TWIX [ 43 ]
24merged6
0.7527
0.6353
0.6890
NightSun [ 36 ]
run16
0.6290
0.6053
0.6169
Graphwise [ 59 ]
9
0.6217
0.5897
0.6053
TUGW [ 27 ]
EXP4SUB1EL
0.7723
0.4437
0.5636
MindGut link
BiomedBERT
0.5738
0.5330
0.5527
GetGut@AAU [ 26 ]
1
0.5515
0.5519
0.5517
Table 15: GutBrainIE performance metrics of each team’s top run beating the baseline for NERD. The best result is in bold, the second-best is underlined (micro-averaged).
Team ID
Run name
Precision
Recall
F1
TWIX [ 43 ]
mre10
0.8996
0.4651
0.6132
NightSun [ 36 ]
run18
0.4459
0.4593
0.4525
SMTE [ 45 ]
R2
0.4476
0.3997
0.4223
GetGut@AAU [ 26 ]
1
0.4546
0.3658
0.4054
GutHub [ 64 ]
1
0.4347
0.3754
0.4029
Graphwise [ 59 ]
BGPT5515GEXP8
0.3517
0.4497
0.3947
Table 16: GutBrainIE performance metrics of each team’s top run beating the baseline for M-RE. The best result is in bold, the second-best is underlined (micro-averaged).
Team ID
Run name
Precision
Recall
F1
TWIX [ 43 ]
11mre12
0.4903
0.2691
0.3475
NightSun [ 36 ]
run18
0.2484
0.2798
0.2632
Graphwise [ 59 ]
BGPT5515GEXP4
0.2092
0.2963
0.2452
GetGut@AAU [ 26 ]
1
0.2123
0.1926
0.2020
TEXA [ 69 ]
1
0.1556
0.1383
0.1464
BASELINE
Atlop-3stage
0.1403
0.1292
0.1345
Table 17: Performance metrics of each team’s top run beating the baseline for C-RE. The best result is in bold, the second-best is underlined (micro-averaged).
Biomedical question answering requires not only accurate extraction of information from scientific literature but also reliable integration of evidence across multiple documents. This study presents a question-type-specific large language model (LLM) framework for BioASQ 14b Task B, designed to improve answer robustness and evidence grounding in biomedical question answering. Rather than applying a single prompting strategy to all questions, the framework selects different inference procedures for yes/no, factoid, and list questions according to their distinct reasoning and evaluation requirements. For yes/no questions, snippet shuffling and self-reflection are used to reduce sensitivity to evidence ordering and improve decision stability. For factoid questions, full-snippet input is combined with chain-of-thought-based in-context learning to support accurate biomedical entity identification. For list questions, a multi-agent architecture is employed, in which evidence extraction, candidate generation, answer verification, and final aggregation are handled collaboratively. Preliminary experiments on BioASQ 13b were used to identify effective inference strategies for each question type, and the resulting framework was subsequently evaluated in the official BioASQ 14b Task B challenge. In the official evaluation, our framework showed competitive performance across multiple batches and achieved first place in the factoid subtask of Batch 4. These results demonstrate the effectiveness of combining question-type-specific inference, ensemble prediction, and agent-based verification for reliable biomedical question answering.
Taeyun Roh, Eunha Lee, Wonjune Jang +3
Department of Computer Science and Engineering, Korea University, Seoul, 02841, Republic of Korea · Department of Mathematics, Myongji University, Yongin, 17058, Republic of Korea · 3AIGEN Sciences, Seoul, 04778, Republic of Korea
We describe our BioASQ Task 14B 2026 system. The work centers on two design decisions: how aggressively to re-retrieve when first-stage retrieval is weak, and how to combine multiple language-model answers. Retrieval unions two parallel pipelines - a hybrid first stage (dense BGE + BM25 + RRF, reaching R@200 = 99.3% on the BioASQ-13b historical archive) and an agent-driven pipeline that decomposes the question over PubMed, Europe PMC, and iCite - with a BGE cross-encoder quality gate flagging weakly-supported questions for selective re-retrieval. On Task 12B 2024 validation, a cost-pragmatic re-retrieval policy beats a skill-strict baseline significantly on list F1 and list precision, at 12% lower re-retrieval cost. Holding prompt and model fixed across val and test 13B (different question sets), list F1 rises by +0.132 absolute on the BioASQ-released gold-input pool, consistent with substantial retrieval-side headroom. For Phase B answering we decompose multi-model ensemble lift into a selection component bounded by the per-question oracle and a fusion component that aggregators can exceed. The decomposition predicts before any experiment that LLM-as-judge wins on selection-dominated metrics (yes/no, multi-reference ROUGE) but is structurally insufficient on the recall component of fusion-friendly metrics (factoid rank-1, list recall). On Task 13B 2025 our synonym-union resolver wins list recall on every head, while GPT-5.5 solo retains the list-F1 lead because the resolver's wider item set costs precision. On the Task 14B 2026 preliminary leaderboard our team places first on the combined-exact aggregate on three of the eight (phase x batch) leaderboards, wins four individual question-type cells, and takes #1 on Phase B b3 ideal.
This work presents DS@GT ARC BioASQ team's work for a biomedical question answering pipeline, integrating multi-source query expansion, neural reranking, retrieval refinement, and OpenBioLLM-assisted answer generation. The system combines PubMed retrieval with fine-tuned MiniLM-based semantic reranking, Reciprocal Rank Fusion (RRF), and feature-based relevance scoring to improve document ranking quality. To address challenging queries with weak retrieval performance, we introduce a conditional weak-question recovery strategy that applies semantic expansion, relationship-aware augmentation, and selective result merging. A post-retrieval pruning stage further removes redundant or low-relevance snippets while preserving evidence coverage for downstream answer generation. Experimental results on BioASQ evaluation batches demonstrate that the proposed recovery and cleanup strategies substantially improve retrieval robustness and MAP@10 performance on difficult question sets. The final system also incorporates output validation and post-processing steps to ensure formatting consistency and submission reliability across BioASQ phases.
Xueying Zhao, Lee Mai, Balaji Anandganesh
Georgia Institute of Technology, North Ave NW, Atlanta, GA 30332