Organizations: Department of Epidemiology, Vanderbilt University, Nashville, TN, USA · Department of Pediatrics, Vanderbilt University Medical Center, Nashville, TN, USA · Department of Biostatistics, Vanderbilt University Medical Center, Nashville, TN, USA · Department of Biomedical Informatics, Vanderbilt University Medical Center, Nashville, TN, USA
Pretrained large language models offer a practical foundation for learning useful behavior from few task-specific examples. We argue that current prompt and context optimization methods underuse the extensive knowledge and reasoning capabilities of trillion-parameter models. These capabilities can make adaptation more sample-efficient, more compute efficient and at no performance loss when organized around how human experts investigate failures. We formalize Policy Iteration with Human Feedback (PIHF), which makes this implicit procedure explicit for LLM agents to execute, and build its automated implementation, PIHF-MCP. Initialized from clinician feedback on rare-disease diagnosis, PIHF-MCP supplies the expert procedure, testing tools, review and persistent inquiry records to develop reusable task policies. Across general reasoning benchmarks (BIG-Bench Extra Hard, HoVer and LiveBench-Math), PIHF-MCP improved performance of the baseline model by 16.9, 22.2 and 4.7 percentage points, respectively. With a matched baseline model, development used about 1/5 of the labelled examples and 4% of the task rollouts reported by a previous SOTA in-context optimizer, making it about 9 times faster and 3 times cheaper at comparable or higher scores. In a low-data rare-disease diagnosis setting, policies developed from previous SOTA prompt optimizers trailed a previously published PIHF-developed system on every held-out cohort (on average 16 percentage points). These findings support a route to more efficient inference-time scaling: PIHF-MCP develops reusable policies from a few examples that improve performance on unseen cases and across models. Because each policy comes from an explicit, recorded investigation, the process also keeps humans in the loop and enables ownership and learning, making it well suited to high-stakes decisions.
Figures & tables
Figure 1: PIHF-MCP turns expert failure investigation into executable learning loops. A , Conventional prompt-optimization rounds repeatedly run, reflect, rewrite and validate. B , PIHF-MCP organizes inquiry around observed failure modes; the Causal Investigator selects tests and revisions, drawing on tools, reviewers and curator memory within the task contract. Clinician-derived investigative practice informs its guidance. The Meta Verifier is the Verifier role described in Methods.
LIRICAL
PhenoPacket Store
Development (n=50)
Held-out (n=320)
Unmapped (n=150)
Mapped (n=150)
Extension (n=573)
System
R@1
R@5
R@1
R@5
R@1
R@5
R@1
R@5
R@1
R@5
Baseline a
42.0
54.0
33.1
47.8
6.0
8.7
40.0
56.7
10.1
13.8
GEPA
62.0
74.0
46.2
60.6
35.3
42.7
60.0
74.7
37.3
46.2
ACE
68.0
74.0
49.1
62.5
28.0
36.7
48.0
68.7
37.7
46.6
PIHF a
66.0
82.0
54.4
71.2
60.7
78.7
64.7
81.3
56.5
75.0
Table 1: PIHF, GEPA and ACE policies developed from the same 50 clinical cases.
Benchmark
Starting policy
PIHF-MCP
Gain
test %
test %, mean ± SD
pp
HoVer (300 test claims)
46.0
68.2 ± 2.2
+22.2
LiveBench-Math (126 test questions)
73.9
78.6 ± 1.0
+4.7
BIG-Bench Extra Hard, mini (177 test questions)
25.4
42.4 ± 2.5
+16.9
Table 2: Performance gains from PIHF-MCP policy development.
Labelled
Task
Development
Test score
Method
examples
rollouts
cost (USD)
time (h)
%, mean ± SD
HoVer (300 test claims)
GEPA
450
7,051 a
34.09 c
16.1 c
51.7 n.r. b
PIHF-MCP
95
259
8.49
1.5
59.9 ± 0.4
PIHF-MCP advantage
4.7 ×
27 ×
4.0 ×
11 ×
+8.2 pp
LiveBench-Math (126 test questions)
Table 3: Development efficiency of PIHF-MCP versus GEPA with the same executor.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Tokens (thousands)
Actor
Model
LM calls
Input
cached
Output
Cost (USD)
HoVer (300 test claims), GPT-4.1-mini executor
Executor
GPT-4.1-mini
740
1,467.4
124.8
448.2
1.27
Causal Investigator
Sol-medium
5
11,246.5
10,799.5
50.0
7.11
Critic
Luna-high
3
44.8
1.8
9.3
0.02
Meta Verifier (pre)
Luna-high
4
117.8
45.1
7.9
0.02
Appendix
Table A1: Development tokens and cost by actor.
Condition
Investigator
Examples
Rollouts
Test %
Gain
dev + val
dev + val
mean ± SD
pp vs. P0
Causal Investigator
BBEH mini (177 test questions), Luna-low executor
Starting policy
25.4
Detailed prompt (Table 2)
Sol-medium
46
134
42.4 ± 2.5
+16.9
Detailed prompt
DeepSeek-high
46
281
41.2
+15.8
Appendix
Table A2: Causal Investigator backbone and guidance on BBEH Mini.
GPT-4.1-mini executor
Luna-low executor
Policy
test %
pp vs. P0
test %
pp vs. P0
A. Across executors
HoVer, original test roster (300 claims)
Starting policy
43.3
46.0
Developed with GPT-4.1-mini
59.9 ± 0.4
+16.6
78.7
+32.7
Developed with Luna-low
56.7
+13.3
68.2 ± 2.2
+22.2
Appendix
Table A3: Policy transfer across executors and test sets.
Large language model (LLM) agents can retrieve memory, call tools, ask clarifying questions, and vary response style, yet adapting these execution decisions to an individual user remains difficult. Fine-tuning a separate LLM is costly or impossible for proprietary systems, while prompts and memory primarily expose user information to the agent rather than adapt its execution decisions from feedback. We formulate personalization of a frozen agent as online learning of a per-user execution policy from scalar feedback observed only for the executed action. We propose FABLE (Factorized Adaptive Bandit Layer for Execution), a lightweight policy layer outside a potentially black-box host agent. FABLE factorizes memory, information-acquisition, and response decisions so feedback updates related choices; filters actions through an externally specified feasible set before exploration; and learns user-specific residual preferences relative to a fixed default-and-cost score via Bayesian contextual Thompson sampling. Under a linear residual-reward model, a calibrated variant inherits an expected-regret bound against the best feasible action. We also characterize preferences unidentifiable under persistent feasibility constraints and provide anytime-valid false-promotion control. Across personalized-reasoning, controlled-feedback, and executable tool-use evaluations, FABLE improves several preference-sensitive behaviors relative to rule-only control while remaining competitive on end-to-end task performance.
Dian Jin, Zhi Zhang, Huichao Li +3
Department of Statistics and Data Science, National University of Singapore · Department of Statistics and Data Science, University of California, Los Angeles · Department of Data Science, Fudan University
Large language models (LLMs) have shown remarkable in-context learning (ICL) capabilities, yet their potential for sequential decision-making remains underexplored. In this paper, we study the ICL capabilities of LLMs in sequential decision-making settings, including Markov Decision Processes (MDPs), Partially Observable MDPs (POMDPs), and Ambiguous POMDPs (APOMDPs). We fine-tune pretrained LLMs to perform few-shot decision-making directly from offline, oracle-labeled trajectories. Our framework enables flexible imitation of policies through supervised fine-tuning (SFT). Theoretically, we focus on linear MDPs and interpret a fine-tuned attention layer as implicitly estimating optimal Q-functions from in-context data. Building on this interpretation, we derive an end-to-end suboptimality bound for the induced policy that separates the in-context estimation error from the training-length bias. Empirically, across synthetic MDP, POMDP, and APOMDP settings, we find that fine-tuned LLMs achieve substantially smaller optimality gaps than in-context-only and random baselines, with especially large gains in longer-horizon, partially observed, and model-ambiguous environments. Together, these results show that supervised fine-tuning provides an effective route to endowing pretrained LLMs with sequential decision-making capabilities from offline data, which is an important advantage in domains such as healthcare where offline data are abundant.
Recent work has shown that large language models (LLMs) can iteratively improve their outputs by incorporating generated samples and their corresponding evaluation scores as in-context examples. Despite these empirical findings, the theoretical foundations underlying this phenomenon remain poorly understood. In this paper, we show that score-conditioned In-Context Learning (ICL) admits a structural correspondence to policy gradient optimization. We first provide a constructive proof that self-attention mechanisms can implement reward-weighted aggregation analogous to the REINFORCE algorithm under specific weight matrix configurations, and discuss the relationship between this construction and the behavior of pretrained transformers. The correspondence is directional in hidden-state space and holds exactly only under the stated simplifying conditions; we quantify its strength empirically. Within our simplified hidden-state model, we furthermore derive an exact upper bound on the distribution shift induced by a bounded attention update, yielding a trust-region-like analogy to KL-constrained policy optimization. We validate our theory through extensive experiments across multiple LLMs, demonstrating that LLMs effectively utilize score information to shift output distributions toward high-scoring exemplars, and that attention weights exhibit a strong correlation with example scores.