cs.CLOct 4, 2026

When Verifiable Counts Depend on Wording: Auditing Wording Robustness in Instruction Following

Authors: Qishi Zhan, Seoyeon Jang, Zihan Dong, Minxuan Hu, Ziheng Chen, Tonghui Qu

Organizations: Marquette University · University of California, San Diego · Georgia Tech · Cornell University · The University of Texas at Austin · Hikvision

Abstract

Verifiable instruction-following benchmarks often express each constraint through one fixed template. We test whether scores remain stable when the operational requirement is unchanged but its wording varies. We introduce WISE, a matched evaluation suite and reporting protocol instantiated on exact word count, keyword inclusion exactly once, and an inclusive 8--12 word range. Across 100 matched tasks, up to thirteen models from seven providers, and repeated generations scored over the complete visible output, wording alone produces substantial compliance shifts. In an avoidance-family panel, five avoidance and exclusion forms fall below the positive baseline, while constructional controls also shift compliance substantially: in the nine-model control panel, compliance is 54.9% for the original positive form, 48.2% for a longer positive form, 36.7% when the target appears later, and 33.8% for AVOID1. A strict JSON-structure probe shows wording sensitivity beyond counting, with a different direction of effect. Effect sizes, failure directions, weakest forms, and model rankings vary across realizations. Under the most disruptive exclusion form, the top-ranked model changes and 24.1% of strictly ordered model pairs reverse. Human validation further shows that unanimous agreement on an exact-count interpretation can coexist with substantially different model behavior. WISE supplements conventional scores with mean and worst-form compliance, wording gaps, failure profiles, and ranking stability.

Figures & tables

Appendix figures & tables24 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Task Competence Is Not Instruction Following: Evaluating Instruction-Conflicting Behavior in Small Language Models

    Jul 21, 2026Mahdiyeh Farajidizaji, Vatsal RainaInstruction-Tuned ModelsInstruction

  2. Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs

    May 19, 2026Carolina Camassa, Derek ShillerInstruction-Tuned ModelsInduction