cs.AIAug 17, 2026

From Answers to Policies: Efficient In-Context Learning System through Emulating Expert Investigation

Authors: Minh-Ha Nguyen, Ngoc-Ngo Quang Tran, Thuy Dung Nguyen, Cathy Shyr

Organizations: Department of Epidemiology, Vanderbilt University, Nashville, TN, USA · Department of Pediatrics, Vanderbilt University Medical Center, Nashville, TN, USA · Department of Biostatistics, Vanderbilt University Medical Center, Nashville, TN, USA · Department of Biomedical Informatics, Vanderbilt University Medical Center, Nashville, TN, USA

Abstract

Pretrained large language models offer a practical foundation for learning useful behavior from few task-specific examples. We argue that current prompt and context optimization methods underuse the extensive knowledge and reasoning capabilities of trillion-parameter models. These capabilities can make adaptation more sample-efficient, more compute efficient and at no performance loss when organized around how human experts investigate failures. We formalize Policy Iteration with Human Feedback (PIHF), which makes this implicit procedure explicit for LLM agents to execute, and build its automated implementation, PIHF-MCP. Initialized from clinician feedback on rare-disease diagnosis, PIHF-MCP supplies the expert procedure, testing tools, review and persistent inquiry records to develop reusable task policies. Across general reasoning benchmarks (BIG-Bench Extra Hard, HoVer and LiveBench-Math), PIHF-MCP improved performance of the baseline model by 16.9, 22.2 and 4.7 percentage points, respectively. With a matched baseline model, development used about 1/5 of the labelled examples and 4% of the task rollouts reported by a previous SOTA in-context optimizer, making it about 9 times faster and 3 times cheaper at comparable or higher scores. In a low-data rare-disease diagnosis setting, policies developed from previous SOTA prompt optimizers trailed a previously published PIHF-developed system on every held-out cohort (on average 16 percentage points). These findings support a route to more efficient inference-time scaling: PIHF-MCP develops reusable policies from a few examples that improve performance on unseen cases and across models. Because each policy comes from an explicit, recorded investigation, the process also keeps humans in the loop and enables ownership and learning, making it well suited to high-stakes decisions.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Personalizing Large Language Model Agents with Small Policy Models

    Jul 31, 2026Dian Jin, Zhi Zhang, Huichao Li +3Large Language Model AgentsSmall Reasoning Model

  2. In-Context Learning as Implicit Policy Gradient

    Jul 25, 2026Masahiro Kaneko, Timothy BaldwinIn-Context LearningIn-Context