UniBuc at SemEval-2024 Task 2: Tailored Prompting with Solar for Clinical NLI
Authors: Marius Micluta-Campeanu, Claudiu Creanga, Ana-Maria Bucur, Ana Sabina Uban, Liviu P. Dinu
Organizations: Interdisciplinary School of Doctoral Studies, ♡HLT Research Center University of Bucharest, Romania · Faculty of Mathematics and Computer Science
This paper describes the approach of the UniBuc team in tackling the SemEval 2024 Task 2: Safe Biomedical Natural Language Inference for Clinical Trials. We used SOLAR Instruct, without any fine-tuning, while focusing on input manipulation and tailored prompting. By customizing prompts for individual CTR sections, in both zero-shot and few-shots settings, we managed to achieve a consistency score of 0.72, ranking 14th in the leaderboard. Our thorough error analysis revealed that our model has a tendency to take shortcuts and rely on simple heuristics, especially when dealing with semantic-preserving changes.
Figures & tables
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Section
F1
Eligibility (Single)
0.6637
Eligibility (Comparison)
0.5274
Intervention (Single)
0.7306
Intervention (Comparison)
0.7751
Results (Single)
0.7141
Results (Comparison)
0.7525
Appendix
Table 2: Initial results for the train set (first 50 examples)
Section
F1
Eligibility (Single)
0.8178
Eligibility (Comparison)
0.8285
Intervention (Single)
0.7678
Intervention (Comparison)
0.6000
Results (Single)
0.7368
Results (Comparison)
0.7749
Appendix
Table 3: Initial results for the development set
Section (CTR type)
Prompt
Eligibility (Single)
Instruction: You are given clinical trial criteria and a statement that may or may not be contradictory. Regarding the inclusion and exclusion criteria, is the statement correct? Respond only with Yes or No. ## Criteria: {premise}. ## Statement: {hypothesis}. ## Response (Yes or No):
Eligibility (Comparison)
Instruction: You are given clinical trial criteria for a primary and a secondary trial, and a statement. Regarding the inclusion and exclusion criteria, is the statement correct for each trial? Respond only with Yes or No. ## Criteria: {premise}. ## Statement: {hypothesis}. ## Response (Yes or No):
Intervention (Single)
Instruction: You are given a CTR and a statement. Can the statement be deduced from the CTR? Focus on the interventions. Respond only with Yes or No. ##CTR: {premise}. ##Statement: {hypothesis}. ##Response (Yes or No):
Intervention (Comparison)
(same prompt template as single CTR for interventions)
Results (Single)
Instruction: You are given the results of a CTR and a statement. Can the statement be deduced from the CTR in terms of number of participants, measures and results? Respond only with Yes or No. ##CTR: {premise}. ##Statement: {hypothesis}. ##Response (Yes or No):
Results (Comparison)
Instruction: You are given the results of a CTR and a statement. Can the statement be deduced from the CTR? Respond only with Yes or No. ##CTR: {premise}. ##Statement: {hypothesis}. ##Response (Yes or No):
Appendix
Table 4: Evaluation templates for each CTR section
Section (CTR type)
Prompt
Eligibility (Single)
Instruction: You are given the eligibility criteria for a clinical trial report. You must summarize the report focusing on inclusion and exclusion criteria. Report: {premise}. Use short sentences. Summary:
Eligibility (Comparison)
(same prompt template as single CTR for eligibility)
Intervention (Single)
Instruction: You are given the intervention information for a clinical trial report. Each report contains 1-2 cohorts, which receive different treatments, or have different characteristics. You must summarize the report focusing on the type, dosage, frequency, and duration of treatments being studied. Report: {premise}. Use short sentences to group by cohort. Summary:
Intervention (Comparison)
Instruction: You are given the intervention information for a clinical trial report. Each report contains 1-2 cohorts, which receive different treatments, or have different characteristics. You must summarize the report focusing on the type, dosage, frequency, and duration of treatments being studied. Report: {premise}. Use short sentences to group by cohort and group by primary trial and secondary trial. Summary:
Results (Single)
Instruction: You are given the results of a CTR and a statement. Extract all the relevant information from the CTR that is related to the statement. Report: {premise}. Statement: {hypothesis}. Answer:
Results (Comparison)
Instruction: You are given the results of two clinical trials. Each trial contains 1-2 cohorts, which receive different treatments, or have different characteristics. You must summarize the report for each trial, focusing on number of participants, outcome measures, units, results. Report: {premise}. Use short sentences and keep all numeric values. Summary:
Appendix
Table 5: Summarization templates for each CTR section. The results section (single CTR) is the only one for which summaries depend on the hypothesis due to lack of time to rerun the summaries for the test set.
Instruction embedding models have become common among state-of-the-art models, however are evaluated using a single prompt per task. The single-point evaluation ignores a main problem of the instruction-based approach namely: sensitivity to the phrasing of the instruction. We present an empirical study of prompt sensitivity across 6 embedding models, 11 datasets, and 15 task-specific prompts per dataset, a total of 990. We show that reported scores misrepresent the distribution of scores over plausible prompts. The default prompt can both systematically understate or overstate performance. Furthermore, we show that the leaderboard ranking is not robust to prompt selection: by choosing prompts favorably, any model in our study can be promoted to first place. Our findings suggest that single-prompt evaluation is insufficient for instruction-tuned embedding models and that benchmarks should incorporate prompt robustness, either by evaluating over multiple prompts or by reporting sensitivity alongside point estimates.
Large language models are increasingly used to summarize clinical trial results for healthcare providers, patients, and payers, but their tendency to hallucinate poses significant risks in this high-stakes context. This study introduces a benchmark evaluation framework for measuring the faithfulness of LLM-generated clinical trial summaries across three stakeholder audiences. The framework consists of 200 stratified trials drawn from the Aggregate Analysis of ClinicalTrials.gov database, evaluated using audience-specific prompt templates and a six-dimension faithfulness annotation schema. Baseline measurements were established for GPT-4o, Claude Sonnet 4.6, and Gemini 2.5 Flash across 1,800 generated summaries scored using a cross-encoder natural language inference (NLI) model. Unsupported Claims was identified as the dominant failure mode across all three models, with a mean annotation score of 1.55 out of three. A knowledge-graph-augmented retrieval system was developed and evaluated against the baseline, producing statistically significant improvements in NLI-based faithfulness scores (entailment +0.0125, faithfulness +0.0130, p < 0.0001). Improvement pathways were model-dependent, with GPT-4o improving primarily through contradiction reduction while Claude Sonnet 4.6 and Gemini 2.5 Flash improved through increased entailment.
Robert Williams
University of Texas at Austin Computer & Data Science Online Austin, Texas, USA
In this paper, we present the RETUYT-INCO participation at the BEA 2026 shared task "Rubric-based Short Answer Scoring for German". Our team participated in track 1 (Unseen answers three-way), track 3 (Unseen answers two-way) and track 4 (Unseen questions two-way). Since these tracks required scoring short student answers using specific rubrics, we looked for ways to handle the changing nature of the task. We created a method called Meta-prompting. In this approach, an LLM creates a custom prompt based on examples from the Train set. This prompt is then used to grade new student answers. Along with this method, we also describe other approaches we used, such as classic machine learning, fine-tuning open-source LLMs, and different prompting techniques. According to the official results, our team placed 6th out of 8 participants in Track 1 with a QWK of 0.729. In Track 3, we secured 4th place out of 9 with a QWK of 0.674, and we also placed 4th out of 8 in Track 4 with a QWK of 0.49.