cs.AISep 12, 2026

IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives

Authors: Kainan Zhou, Zhaoyi Li, Janet Sung, Gangzhen Qian, Hang Xiao

Organizations: Google LLC Mountain View, CA, USA

Abstract

An external record may contain a procedure to apply or text to read, depending on the user's request. IBBench-Light tests both uses against the same record. Twelve semantic bases yield 144 matched pairs per model; four quantized instruction models produced 1,152 archived greedy responses. Paired exact-contract accuracy (PECA) requires both members to satisfy their output contracts. Qwen succeeds on 132 execute and 109 process prompts, but only 97 complete pairs, showing what marginal averages omit. We audit literal-target exposure and case normalization, then add 1,722 logged CPU generations to test directive-absent controls, twelve additional semantic bases, within-base wording changes, and generation stopping. In the pinned Phi rerun, changing the end-of-sequence (EOS) set changes exact paired success from 0/144 to 62/144. A bounded IHEval comparison uses the same SmolLM2 checkpoint and output budget while preserving its published instruction roles and scorer. The benchmark measures conditional task and output-contract success. Its task margins and paired count need to be read together with the stopping policy.

Figures & tables

Explore similar work

CardsList
  1. Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets

    Jul 27, 2026Zongyou Yang, Yinghan HouMulti-Dimensional EvaluationMath Problems

  2. ContractEval: Query-Conditioned Execution Matching for Procedural Instruction Conformance

    Sep 10, 2026Praphul Singh, Shanu Kumar, Akshat Agarwal +1ContractsLarge Language Model Judges