cs.LGOct 5, 2026

Dynamic Budget Allocation for LLM Evaluation under Hard Resource Constraints

Authors: Shai Feldman, Yaniv Romano

Organizations: Department of Computer Science, Technion IIT, Israel · Department of Electrical and Computer Engineering, Technion IIT, Israel

Abstract

We evaluate large language models (LLMs) in multi-turn interactions through their time-to-event: the number of interaction steps required to produce an event of interest, such as a successful jailbreak or agentic task completion. Under limited compute, interactions may be terminated before the event occurs, so that event times are only partially observed (censored). Existing allocation methods for calibrating time-to-event bounds satisfy the budget only in expectation and can exceed the available budget on a particular evaluation run. Enforcing a hard constraint is particularly challenging as the cost of a trajectory is initially unknown. We introduce Hard-budget Allocation with Reflow for Predictive calibration (HARP), a budget allocation that satisfies hard resource constraints and adaptively reallocates unused budget. We show how to use HARP to construct lower predictive bounds (LPBs) on the time-to-event and to estimate evaluation metrics such as the jailbreak rate on a fixed benchmark. Although HARP induces dependence in acquisition decisions across different trajectories, we prove that HARP never exceeds the target budget, that its LPBs have finite-sample coverage guarantees, and that its metric estimates are unbiased. Experiments on agentic task success, LLM jailbreaks, toxic content generation, and RAG hallucinations show that HARP achieves coverage close to the nominal level with low variance, while never exceeding the given budget.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation

    May 7, 2026Shai Feldman, Yaniv RomanoLarge Language Model JailbreaksJailbreak Attacks

  2. Adaptive Test-Time Compute Allocation for Reasoning LLMs via Constrained Policy Optimization

    Apr 16, 2026Zhiyuan Zhai, Bingcong Li, Bingnan Xiao +2Test-Time ScalingLLM Reasoning Strategies

  3. Cost-Aware Multi-Objective Bandits: Theory and Application to Budgeted LLM Configuration Evaluation

    Aug 5, 2026Bo Xue, Zhi Hong, Jiayi Li +3Cost-AwareLarge Language Model Evaluation