Few-Shot Prompting

Momentum

3 papers in the last four weeks, with none the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 34

Dec 17, 2025cs.LG

Improving Fairness of Large Language Model-Based ICU Mortality Prediction via Case-Based Prompting

Accurately predicting mortality risk in intensive care unit (ICU) patients is critical for clinical decision-making. Large language models (LLMs) are increasingly explored for clinical prediction using structured medical data, but their outputs may exhibit demographic disparities. Mitigating such disparities without degrading predictive performance remains challenging. We systematically investigate demographic bias in LLM-based ICU mortality prediction and propose Case-Based Prompting (CAP), a training-free framework designed to improve the empirical balance between predictive performance and subgroup fairness. CAP retrieves clinically similar historical misprediction and demographic-sensitive cases with known outcomes and incorporates them as case-level contextual evidence. We further evaluate model discrimination, subgroup fairness, and feature-dependence consistency. On MIMIC-IV, CAP improved AUROC from 0.806 to 0.873 and AUPRC from 0.497 to 0.694 compared with the Baseline Prompt. CAP also reduced several subgroup disparities, particularly for sex and White-Black comparisons, although age-related disparities remained non-negligible. Controlled demographic-information ablations showed that removing demographic information did not uniformly improve both predictive performance and subgroup fairness, indicating distinct performance-fairness patterns rather than a uniform benefit from demographic suppression. Feature-dependence analysis showed generally consistent expressed clinical factor patterns across demographic subgroups. These findings suggest that CAP is a promising inference-time strategy for LLM-based clinical prediction under the evaluated retrospective setting. Further validation across additional LLMs, external datasets, and prospective clinical scenarios is required before deployment.
Jun 8, 2025cs.LG

Training-free LLM Verification via Recycling Few-shot Examples

Although large language models (LLMs) have achieved remarkable performance, the inherent stochasticity of their reasoning processes and varying conclusions present significant challenges. Majority voting or Best-of-N with external verifiers has been explored to mitigate this, but these approaches are limited in applicability or require additional training. To address this problem, we propose a novel framework that Recycles Few-shot examples to verify LLM outputs (ReFeri). Our key idea is to utilize the given few-shot examples not only to generate outputs, but also to evaluate the candidate outputs. Specifically, ReFeri combines a forward confidence score with a backward reconstruction penalty to select candidates that follow few-shot guidance while avoiding demonstration-specific overfitting. Experiments with three different LLMs across seven diverse tasks demonstrate that our framework significantly improves the accuracy of LLMs---achieving an average relative gain of 8.2%---through effective response selection.
Mar 7, 2024cs.CL

Few shot chain-of-thought driven reasoning to prompt LLMs for open ended medical question answering

In this paper, we propose a modified version of the MedQA-USMLE dataset, named MEDQA-OPEN, which contains open-ended medical questions without options to mimic clinical scenarios, along with clinician-approved reasoned answers. Additionally, we implement a prompt driven by Chain of Thought (CoT) reasoning, CLINICR, to mirror the prospective process of incremental reasoning, reaching a correct response to medical questions. We empirically demonstrate how CLINICR outperforms the state-of-the-art 5-shot CoT-based prompt (Liévin et al., 2022). We also present an approach that mirrors real-life clinical practice by first exploring multiple differential diagnoses through MCQ-CLINICR and subsequently narrowing down to a final diagnosis using MCQ-ELIMINATIVE. Finally, emphasizing the importance of response verification in medical settings, we utilize a reward model mechanism, replacing the elimination process performed by MCQ-ELIMINATIVE.
Date pendingcs.CL

Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization

The growing complexity of content moderation policies presents a critical challenge for their consistent operationalization. While foundation models possess the basic capabilities needed to confront this challenge, whether they can reliably moderate online content remains an unanswered question. In this paper, we systematically compare two competing paradigms for Vision-Language Model (VLM) guidance: an instruction-driven approach where models reason from policy precepts, and an example-driven approach where they generalize from prior precedents. We ground this investigation in ModerationBench, a new benchmark of 4,000 manually annotated, in-the-wild posts from the Bluesky platform. Our experiments reveal that foundation models can substantially outperform Bluesky's deployed moderation system, nearly tripling its F1F_1 score (0.60 vs. 0.22) on Random Posts in the benchmark, with both instruction- and example-driven paradigms achieving comparable peak effectiveness. Our findings thus chart a path toward reliable and adaptable policy operationalization at scale.