Organizations: Department of Epidemiology, Vanderbilt University, Nashville, TN, USA · Department of Pediatrics, Vanderbilt University Medical Center, Nashville, TN, USA · Department of Biostatistics, Vanderbilt University Medical Center, Nashville, TN, USA · Department of Biomedical Informatics, Vanderbilt University Medical Center, Nashville, TN, USA
Pretrained large language models offer a practical foundation for learning useful behavior from few task-specific examples. We argue that current prompt and context optimization methods underuse the extensive knowledge and reasoning capabilities of trillion-parameter models. These capabilities can make adaptation more sample-efficient, more compute efficient and at no performance loss when organized around how human experts investigate failures. We formalize Policy Iteration with Human Feedback (PIHF), which makes this implicit procedure explicit for LLM agents to execute, and build its automated implementation, PIHF-MCP. Initialized from clinician feedback on rare-disease diagnosis, PIHF-MCP supplies the expert procedure, testing tools, review and persistent inquiry records to develop reusable task policies. Across general reasoning benchmarks (BIG-Bench Extra Hard, HoVer and LiveBench-Math), PIHF-MCP improved performance of the baseline model by 16.9, 22.2 and 4.7 percentage points, respectively. With a matched baseline model, development used about 1/5 of the labelled examples and 4% of the task rollouts reported by a previous SOTA in-context optimizer, making it about 9 times faster and 3 times cheaper at comparable or higher scores. In a low-data rare-disease diagnosis setting, policies developed from previous SOTA prompt optimizers trailed a previously published PIHF-developed system on every held-out cohort (on average 16 percentage points). These findings support a route to more efficient inference-time scaling: PIHF-MCP develops reusable policies from a few examples that improve performance on unseen cases and across models. Because each policy comes from an explicit, recorded investigation, the process also keeps humans in the loop and enables ownership and learning, making it well suited to high-stakes decisions.
Figures & tables
Figure 1: PIHF-MCP turns expert failure investigation into executable learning loops. A , Conventional prompt-optimization rounds repeatedly run, reflect, rewrite and validate. B , PIHF-MCP organizes inquiry around observed failure modes; the Causal Investigator selects tests and revisions, drawing on tools, reviewers and curator memory within the task contract. Clinician-derived investigative practice informs its guidance. The Meta Verifier is the Verifier role described in Methods.
LIRICAL
PhenoPacket Store
Development (n=50)
Held-out (n=320)
Unmapped (n=150)
Mapped (n=150)
Extension (n=573)
System
R@1
R@5
R@1
R@5
R@1
R@5
R@1
R@5
R@1
R@5
Baseline a
42.0
54.0
33.1
47.8
6.0
8.7
40.0
56.7
10.1
13.8
GEPA
62.0
74.0
46.2
60.6
35.3
42.7
60.0
74.7
37.3
46.2
ACE
68.0
74.0
49.1
62.5
28.0
36.7
48.0
68.7
37.7
46.6
PIHF a
66.0
82.0
54.4
71.2
60.7
78.7
64.7
81.3
56.5
75.0
Table 1: PIHF, GEPA and ACE policies developed from the same 50 clinical cases.
Benchmark
Starting policy
PIHF-MCP
Gain
test %
test %, mean ± SD
pp
HoVer (300 test claims)
46.0
68.2 ± 2.2
+22.2
LiveBench-Math (126 test questions)
73.9
78.6 ± 1.0
+4.7
BIG-Bench Extra Hard, mini (177 test questions)
25.4
42.4 ± 2.5
+16.9
Table 2: Performance gains from PIHF-MCP policy development.
Labelled
Task
Development
Test score
Method
examples
rollouts
cost (USD)
time (h)
%, mean ± SD
HoVer (300 test claims)
GEPA
450
7,051 a
34.09 c
16.1 c
51.7 n.r. b
PIHF-MCP
95
259
8.49
1.5
59.9 ± 0.4
PIHF-MCP advantage
4.7 ×
27 ×
4.0 ×
11 ×
+8.2 pp
LiveBench-Math (126 test questions)
Table 3: Development efficiency of PIHF-MCP versus GEPA with the same executor.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Tokens (thousands)
Actor
Model
LM calls
Input
cached
Output
Cost (USD)
HoVer (300 test claims), GPT-4.1-mini executor
Executor
GPT-4.1-mini
740
1,467.4
124.8
448.2
1.27
Causal Investigator
Sol-medium
5
11,246.5
10,799.5
50.0
7.11
Critic
Luna-high
3
44.8
1.8
9.3
0.02
Meta Verifier (pre)
Luna-high
4
117.8
45.1
7.9
0.02
Appendix
Table A1: Development tokens and cost by actor.
Condition
Investigator
Examples
Rollouts
Test %
Gain
dev + val
dev + val
mean ± SD
pp vs. P0
Causal Investigator
BBEH mini (177 test questions), Luna-low executor
Starting policy
25.4
Detailed prompt (Table 2)
Sol-medium
46
134
42.4 ± 2.5
+16.9
Detailed prompt
DeepSeek-high
46
281
41.2
+15.8
Appendix
Table A2: Causal Investigator backbone and guidance on BBEH Mini.
GPT-4.1-mini executor
Luna-low executor
Policy
test %
pp vs. P0
test %
pp vs. P0
A. Across executors
HoVer, original test roster (300 claims)
Starting policy
43.3
46.0
Developed with GPT-4.1-mini
59.9 ± 0.4
+16.6
78.7
+32.7
Developed with Luna-low
56.7
+13.3
68.2 ± 2.2
+22.2
Appendix
Table A3: Policy transfer across executors and test sets.
Department of Statistics and Data Science, National University of Singapore · Department of Statistics and Data Science, University of California, Los Angeles · Department of Data Science, Fudan University