cs.LGApr 15, 2026

PAC-CF: Calibrating Irreversible Frontier Pruning in LLM-Guided Search

Authors: Tianhao Qian, Jiayu Chen, Zhenyu Sun, Lixu Wang

Abstract

LLM-guided search explores multiple candidate trajectories, but at substantial test-time cost. Pruning low-scoring frontier candidates can control this cost, yet it also turns potentially biased evaluator scores into irreversible decisions: systematic ranking errors can persist under repeated scoring and remove useful branches. We propose Probably Approximately Correct Conformal Filtering (PAC-CF). Its fixed-frontier analysis formulates elimination as an (ε,δ)(\varepsilon,δ)-PAC problem under bounded evaluator bias; its operational rule separately calibrates a score-gap threshold on held-out tasks by running the original controller without PAC-CF and using post-search verifier labels to measure the deficit of solution-preserving candidates relative to the frontier leader. Conditional on exchangeable native-controller tasks with nonempty protected exposure, conformal calibration gives finite-sample coverage for retaining at least one verifier-defined valid continuation at every protected frontier on the native trajectory. At deployment, PAC-CF removes only candidates whose gap from the highest frontier score exceeds the frozen threshold. We evaluate PAC-CF across three domains, five controllers, and four request budgets from B100 to B500. In the cross-domain/controller macro averages, the point estimates for all three workload measures are lower at every budget; the paired-bootstrap 95% confidence interval for utility excludes zero at B100 and B200. For pruning-aware ToolTree, the full-test-set cross-domain utility difference is +4.38+4.38 points at each tested budget; on the natural-termination sensitivity cohort, physical requests decrease by 18.9418.94--18.95%18.95\% and end-to-end token usage by 23.5723.57--23.76%23.76\%.

Figures & tables

Appendix figures & tables19 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility

    Jul 29, 2026Yansen Zhang, Yilu Liu, Tianyu Liu +6Cost-AwareFrontiers

  2. Agentic Search for Counterfactual Recourse under Fixed LLM Budgets

    Jun 7, 2026Yasuo TabeiAgentic SearchToken Budget Allocation

  3. Conformal Cascade: Distribution-Free Accuracy Guarantees for Multi-Tier LLM Inference

    Jul 27, 2026Yifan Dou, Shikan Lian, Shibo LiLLM Inference OptimizationConformal Selection