Organizations: Marquette University · University of California, San Diego · Georgia Tech · Cornell University · The University of Texas at Austin · Hikvision
Verifiable instruction-following benchmarks often express each constraint through one fixed template. We test whether scores remain stable when the operational requirement is unchanged but its wording varies. We introduce WISE, a matched evaluation suite and reporting protocol instantiated on exact word count, keyword inclusion exactly once, and an inclusive 8--12 word range. Across 100 matched tasks, up to thirteen models from seven providers, and repeated generations scored over the complete visible output, wording alone produces substantial compliance shifts. In an avoidance-family panel, five avoidance and exclusion forms fall below the positive baseline, while constructional controls also shift compliance substantially: in the nine-model control panel, compliance is 54.9% for the original positive form, 48.2% for a longer positive form, 36.7% when the target appears later, and 33.8% for AVOID1. A strict JSON-structure probe shows wording sensitivity beyond counting, with a different direction of effect. Effect sizes, failure directions, weakest forms, and model rankings vary across realizations. Under the most disruptive exclusion form, the top-ranked model changes and 24.1% of strictly ordered model pairs reverse. Human validation further shows that unanimous agreement on an exact-count interpretation can coexist with substantially different model behavior. WISE supplements conventional scores with mean and worst-form compliance, wording gaps, failure profiles, and ranking stability.
Figures & tables
Figure 1: A matched example of wording-sensitive exact-count compliance. The model, base task, target, checker, and decoding configuration are fixed; only the constraint wording changes.
Form
Template
POS
Answer using exactly N words.
NEG
Do not use more or fewer than N words.
NEVER
Never use more or fewer than N words.
AVOID1
Avoid using any number of words other than N .
AVOID2
Avoid giving an answer that is not exactly N words.
CNEG
Do not fail to answer using exactly N words.
Table 1: Wording forms in the original WISE exact-count panel. All forms require the complete visible output to contain exactly N words.
Figure 2: Strict compliance across three constraints for the six shared non-Gemini models. AVOID is lowest throughout. NEG has larger negative effects for exact count and keyword inclusion than for the 8–12 range, whose interval includes zero. Exact-count AVOID denotes AVOID1.
Form
Compliance
Diff. vs. POS
95% CI
Mean abs. error
POS
49.5
−
−
0.71
NEG
39.0
-10.4
[-17.1, -4.4]
1.00
NEVER
40.8
-8.6
[-14.6, -2.6]
0.91
AVOID1
28.9
-20.5
[-29.0, -12.9]
2.96
AVOID2
46.9
-2.6
[-6.3, 3.2]
0.83
CNEG
50.1
+0.7
[-2.2, 4.5]
0.67
Table 2: Original exact-count results for the eleven non-Gemini models. Values are percentages; intervals are 95% provider–model–item bootstrap intervals against POS.
Form
Compliance
Diff. vs. POS
95% CI
POS
54.9
−
−
POS_LONG
48.2
-6.7
[-9.2, -4.3]
POS_END
36.7
-18.2
[-21.3, -15.1]
AVOID1
33.8
-21.1
[-24.4, -17.9]
Table 3: Constructional-control analysis for nine retained models. POS_LONG is a longer affirmative realization; POS_END is an affirmative realization with the target later in the sentence.
Form
Compliance
Diff. vs. POS
95% CI
Under
Over
POS
49.2
−
−
31.5
19.3
AVOID1
27.9
-21.3
[-29.1, -14.0]
33.7
38.4
AVOID3
34.0
-15.2
[-23.2, -9.6]
52.1
13.9
AVOID4
25.2
-23.9
[-32.3, -16.6]
40.1
34.6
EXCLUDE1
18.7
-30.5
[-38.9, -22.6]
44.0
37.3
EXCLUDE2
36.9
-12.2
[-18.1, -6.1]
49.2
13.8
Table 4: Avoidance-family panel for the eleven non-Gemini models. Values are percentages; intervals are 95% provider–model–item bootstrap intervals against POS.
Form
τb
95% CI
Reversal
95% CI
AVOID1
0.624
[0.367, 0.722]
18.5
[13.2, 31.5]
AVOID3
0.844
[0.587, 0.891]
7.4
[5.5, 20.4]
AVOID4
0.673
[0.440, 0.782]
16.4
[10.9, 27.8]
EXCLUDE1
0.514
[0.273, 0.648]
24.1
[17.0, 36.4]
EXCLUDE2
0.745
[0.514, 0.855]
12.7
[7.3, 24.1]
Table 5: Ranking stability relative to POS in the avoidance-family panel. Reversal is the percentage of model pairs whose strict ordering changes. Intervals are task-bootstrap 95% intervals. The corresponding top-three Jaccard overlaps are 1.00, 0.50, 0.50, 0.20, and 0.50; maximum rank shifts are 5, 2, 4, 4, and 3.
Figure 3: Model rankings across wording forms in the avoidance-family panel. Rows follow the POS ordering. Each cell reports average rank, with strict compliance shown below. Ties receive average statistical ranks. Wording changes both absolute compliance and comparative model ordering.
Form
Compliance
Diff. vs. POS
Parse fail
POS
13.1
−
86.9
NEG
16.1
+3.1
83.9
AVOID
20.6
+7.6
79.4
Table 6: Strict JSON-structure probe across six models. Compliance requires the complete visible output to be valid JSON with exactly the required key set. Values are percentages.
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
Provider
Paper name
API model ID
Accessed
Anthropic
Claude Haiku 4.5
claude-haiku-4-5
June 12, 2026
Anthropic
Claude Sonnet 4.6
claude-sonnet-4-6
June 12, 2026
Anthropic
Claude Opus 4.8
claude-opus-4-8
June 20, 2026
OpenAI
GPT-4o mini
gpt-4o-mini
June 12, 2026
OpenAI
GPT-4.1 mini
gpt-4.1-mini
June 12, 2026
OpenAI
GPT-5.5
gpt-5.5-2026-04-23
June 20, 2026
Appendix
Table 7: Evaluated models and API identifiers. Dynamic aliases are reported exactly as called. The exact-count study uses all thirteen models; the keyword and range extensions use the seven models available when those runs were executed.
Model
Form
Target
POS output
Count
Contrast output
Claude Haiku 4.5
AVOID1
5
Jupiter is the largest planet.
5
Jupiter is the largest planet here. (6, over)
DeepSeek Chat
AVOID4
5
Jupiter is the largest planet.
5
Jupiter (1, under)
GPT-5.5
EXCLUDE1
5
Gold’s chemical symbol is Au.
5
Au (1, under)
Appendix
Table 8: Matched avoidance-family examples in which POS is compliant but an alternative wording fails for the same model, item, target, and repetition. The examples include both over-target and under-target failures.
Model
Form
Target
Count
Output
GPT-4o mini
AVOID1
5
6
The capital of France is Paris.
GPT-4o mini
AVOID4
5
6
The capital of France is Paris.
Claude Sonnet 4.6
EXCLUDE1
5
4
William Shakespeare wrote it.
Claude Sonnet 4.6
EXCLUDE1
8
7
Vaccines train immunity, stopping transmission between people.
GPT-4.1 mini
EXCLUDE2
5
4
William Shakespeare wrote Hamlet.
Appendix
Table 9: Representative failures from the avoidance-family panel under strict entire-output whitespace counting.
Figure 4: Distribution of observed word count minus the requested target in the avoidance-family panel for the eleven primary non-Gemini models. AVOID3 and EXCLUDE2 are concentrated below the target, while AVOID1, AVOID4, and especially EXCLUDE1 exhibit broader bidirectional errors. Boxes show the interquartile range, whiskers extend to 1.5 times the interquartile range, and outliers are not displayed.
Figure 5: POS-to-worst-form compliance gaps for the eleven primary non-Gemini models in the avoidance-family panel. Circles show POS compliance, diamonds show each model’s weakest avoidance or exclusion form, and annotations report the corresponding decrease.
Figure 6: Under-target, exact, and over-target rates in the avoidance-family panel for the eleven non-Gemini models.
Figure 7: Original-form compliance contrasts against POS for the eleven non-Gemini models. Error bars are provider–model–item bootstrap 95% confidence intervals.
Figure 8: Avoidance-family compliance contrasts against POS for the eleven non-Gemini models.
Family
Contrast
Diff.
95% CI
Raw p
Holm p
Original exact-count
NEG–POS
-10.4
[-17.1, -4.4]
<.001
<.001
NEVER–POS
-8.6
[-14.6, -2.6]
.006
.018
AVOID1–POS
-20.5
[-29.0, -12.9]
<.001
<.001
AVOID2–POS
-2.6
[-6.3, 3.2]
.357
.714
CNEG–POS
+0.7
[-2.2, 4.5]
.689
.714
Avoidance family
AVOID1–POS
-21.3
[-29.1, -14.0]
<.001
<.001
Appendix
Table 10: Multiplicity analysis for the primary wording contrasts. Differences are percentage points relative to POS. Raw two-sided bootstrap p -values and 95% confidence intervals are computed from the same bootstrap draws used in the primary analysis. Holm correction is applied separately within each experimental family.
Panel
POS
Cmean
Cworst
Gwording
Original exact count
49.5
42.5
28.9
21.2
Avoidance family
49.2
32.0
18.7
30.5
Constructional controls
54.9
43.4
33.8
21.1
Keyword exactly once
77.4
59.0
44.3
33.2
Inclusive 8–12 range
89.4
85.3
80.6
8.8
JSON structure
13.1
16.6
13.1
7.6
Appendix
Table 11: WISE audit summaries across experimental panels. Values are percentages over each panel’s primary model set. Gaps are computed from unrounded rates. For JSON, POS is the original positive form even though it is not the best-performing form.
Experiment
Target
Exact
Mean words
Mean error
Original
5
30.4
3.66
-1.34
Original
8
0.0
2.85
-5.15
Original
10
0.0
2.89
-7.11
Original
12
0.0
2.93
-9.07
Family
5
30.9
3.56
-1.44
Family
8
0.0
2.76
-5.24
Appendix
Table 12: Gemini 2.5 Flash saved-output floor analysis. Exact is strict compliance in percent; mean error is observed word count minus requested target.
Form
Compliance
Diff.
95% CI
POS
77.4
−
−
NEG
55.3
-22.1
[-25.8, -18.2]
AVOID
44.3
-33.2
[-37.3, -29.1]
Appendix
Table 13: Primary keyword-exactly-once results for the six shared non-Gemini models. Values are percentages; differences and item-cluster bootstrap 95% confidence intervals are relative to POS.
Form
Compliance
Diff.
95% CI
POS
89.4
−
−
NEG
85.8
-3.6
[-10.6, 1.2]
AVOID
80.6
-8.8
[-16.3, -3.2]
Appendix
Table 14: Primary inclusive 8–12 word range results for the six shared non-Gemini models. Values are percentages; differences and model–item bootstrap 95% confidence intervals are relative to POS.
Form
Comp.
Diff.
95% CI
Zero
Multi.
POS
69.7
−
−
22.8
7.5
NEG
49.2
-20.5
[-24.0, -16.7]
47.5
3.2
AVOID
38.7
-31.1
[-34.8, -27.3]
58.9
2.5
Appendix
Table 15: Keyword-exactly-once robustness across all seven models, including Gemini 2.5 Flash. Values are percentages. Diff. and confidence intervals are relative to POS; Zero and Multi. denote zero and multiple keyword occurrences.
Form
Comp.
Diff.
95% CI
Below
Above
POS
76.6
−
−
17.2
6.2
NEG
73.6
-3.0
[-9.4, 1.0]
17.1
9.3
AVOID
69.1
-7.5
[-14.3, -2.2]
22.0
8.8
Appendix
Table 16: Inclusive 8–12 word range robustness across all seven models, including Gemini 2.5 Flash. Values are percentages; intervals use the model–item bootstrap.
Model
POS
POS_LONG
POS_END
AVOID1
Claude Haiku 4.5
49.3
34.7
21.7
5.0
Claude Opus 4.8
80.7
64.3
35.7
83.0
Claude Sonnet 4.6
59.3
63.0
60.0
44.7
DeepSeek Chat
49.3
34.3
16.0
26.3
DeepSeek V4 Pro
59.3
58.7
46.0
28.7
GPT-4.1 mini
53.7
53.0
47.7
39.3
Appendix
Table 17: Per-model compliance in the constructional-control panel. Values are percentages.
Model
POS
NEG
AVOID
Claude Haiku 4.5
0.0
0.0
0.0
Claude Sonnet 4.6
0.0
0.0
0.0
DeepSeek Chat
70.7
62.0
62.3
GPT-4.1 mini
0.7
6.0
5.3
GPT-4o mini
7.0
28.7
55.0
Mistral Small
0.0
0.0
1.0
Appendix
Table 18: Strict JSON-structure compliance by model and wording form. Values are percentages.
Model
POS
NEG
AVOID
Gap
Claude Haiku 4.5
73.7
54.3
49.3
24.3
Claude Sonnet 4.6
100.0
100.0
98.7
1.3
GPT-4o mini
97.0
95.3
88.3
8.7
GPT-4.1 mini
97.0
97.3
93.3
4.0
Gemini 2.5 Flash
0.0
0.0
0.7
0.7
DeepSeek Chat
92.0
90.3
89.0
3.0
Appendix
Table 19: Strict 8–12 word range compliance by model and wording form. Gap is the best–worst difference in percentage points.
Model
POS
NEG
AVOID
Claude Haiku 4.5
78.7
77.7
77.7
Claude Sonnet 4.6
88.0
88.3
91.0
GPT-4o mini
70.3
12.3
9.3
GPT-4.1 mini
91.7
79.7
38.3
Gemini 2.5 Flash
23.3
12.7
5.0
DeepSeek Chat
73.7
44.3
41.0
Appendix
Table 20: Keyword-exactly-once strict compliance by model and wording form. Values are percentages. These are the seven models included when the keyword experiment was run.
Model
POS
NEG
NEVER
AVOID1
AVOID2
CNEG
Claude Haiku 4.5
54.7
38.3
38.7
4.7
54.0
52.7
Claude Sonnet 4.6
58.3
60.0
67.0
41.7
46.7
60.7
Claude Opus 4.8
82.3
78.3
79.3
79.0
81.3
80.3
GPT-4o mini
53.7
38.3
38.3
21.0
48.0
54.7
GPT-4.1 mini
55.3
39.0
39.3
44.0
54.3
60.0
GPT-5.5
62.7
67.0
64.7
43.0
52.7
56.7
Appendix
Table 21: Original six-form strict compliance by model and form. Values are percentages.
Model
POS
AVOID1
AVOID3
AVOID4
EXCLUDE1
EXCLUDE2
Claude Haiku 4.5
53.7
4.0
36.3
4.0
7.7
36.0
Claude Sonnet 4.6
56.3
40.0
57.7
50.3
41.7
54.3
Claude Opus 4.8
80.7
78.3
67.7
59.7
24.0
67.0
GPT-4o mini
55.0
21.7
40.0
27.3
23.3
27.7
GPT-4.1 mini
57.3
42.7
44.3
41.0
30.3
37.3
GPT-5.5
63.7
42.7
57.7
54.3
32.3
62.3
Appendix
Table 22: Avoidance-family strict compliance by model and form. Values are percentages.
Experiment
Contrast
Method
95% CI
Diff.
Original
AVOID1 − POS
Provider–model–item
[-29.0, -12.9]
-20.5
Original
AVOID1 − POS
Model–item
[-28.7, -13.6]
-20.5
Original
AVOID1 − POS
Equal-model item
[-23.1, -17.8]
-20.5
Original
NEG − POS
Provider–model–item
[-17.1, -4.4]
-10.4
Original
NEG − POS
Model–item
[-16.0, -4.6]
-10.4
Original
NEG − POS
Equal-model item
[-12.9, -7.9]
-10.4
Appendix
Table 23: Bootstrap sensitivity analyses for key non-Gemini contrasts. Differences and confidence intervals are percentage points.
Form
A1
A2
A3
A4
A5
Exact- N
POS
1.00
1.00
1.00
1.00
1.00
100
AVOID1
1.00
0.75
1.00
0.75
1.00
90
AVOID3
1.00
1.00
0.75
1.00
1.00
95
AVOID4
1.00
1.00
0.75
0.75
1.00
90
EXCLUDE1
1.00
0.75
1.00
0.75
1.00
90
EXCLUDE2
1.00
1.00
1.00
1.00
1.00
100
Appendix
Table 24: Annotator-level exact-count judgment rates by wording form. Each annotator entry is the proportion of exact- N judgments across target counts of 5, 8, 10, and 12 words. The final column aggregates all twenty judgments for each form.
Measure
Statistic
Estimate
95% CI
Exact- N judgment
Fleiss’ κ
0.54
[0.34, 0.72]
Clarity
ICC(2, k )
0.79
[0.62, 0.90]
Naturalness
ICC(2, k )
0.75
[0.56, 0.88]
Ambiguity
ICC(2, k )
0.68
[0.45, 0.84]
Complexity
ICC(2, k )
0.64
[0.40, 0.81]
Appendix
Table 25: Inter-annotator agreement across all 24 wording and target combinations. Fleiss’ κ is used for the binary exact- N judgment. Continuous ratings use ICC(2, k ), a two-way random-effects, absolute-agreement, average-measures specification. Confidence intervals are obtained from 2000 bootstrap replicates that resample the 24 items while retaining all five annotators for each sampled item.
Instruction tuning is meant to make language models follow user requests, yet it is unclear whether small models comply when an instruction conflicts with their usual task behavior. We study this across three tasks - multiple-choice question answering (MCQA), sentiment classification, and mathematical question answering - by pairing a standard instruction with a conflicting non-standard one (select an incorrect option, output the opposite sentiment, or return twice the answer). This cross-task design allows us to test whether resistance to conflicting instructions is tied to specific task characteristics or reflects a broader behavioral tendency. As all predictions are scored against the original ground truth, a model that ignores the non-standard instruction still appears accurate. Using standard accuracy, non-standard accuracy, and an Instruction-Following Failure Rate (IFFR), we evaluate instruction-tuned Qwen models across sizes. Both standard accuracy and instruction following generally improve with scale, although the pattern is not consistent across all tasks and datasets. Small models stay competent yet routinely ignore the non-standard instruction, while larger models show a clear gap between the two settings. These findings suggest that gains in task capability do not automatically provide reliable control over model behavior. Task competence and instruction following are therefore distinct abilities, and reporting only standard accuracy hides instruction-following failures.
Mahdiyeh Farajidizaji, Vatsal Raina
Khajeh Nasir Toosi University of Technology · Apta AI, Spark AI Research
Language models are trained to follow instructions, but they are also powerful pattern completers. What happens when these two objectives conflict? We construct conversations in which a user instruction to behave in a target way T (e.g., always output a specific token, answer in a particular language, or adopt a persona) is opposed by N hardcoded assistant turns demonstrating a competing pattern P. We then measure instruction-following (IF) rates in this setting, across 13 models and 16 different instructions, for up to 50 turns. Average instruction-following rates range from 1% to 99% across models, largely uncorrelated with standard capability benchmarks. The transition from instruction-following to pattern-following is universal but highly model-dependent. Robustness is modulated both by instruction content, with models resisting induction longer when instructions align with their trained value priors, and by output format, with diverse multi-token responses proving substantially more resistant than single-token outputs. Chain-of-thought reasoning improves robustness but does not eliminate susceptibility, and can produce dissociation between correct deliberation and incorrect output. When asked to predict their behavior in this setting, models achieve 83.5% accuracy on average but systematically underestimate their own resistance to induction pressure. These results suggest that instruction-following remains brittle under induction pressure even for otherwise capable models, and that output diversity, rather than semantic engagement with the input, is the primary factor predicting robustness.
Large reasoning models (LRMs) often improve math and coding performance, but their effect on instruction following is unclear. We study IFEval with Qwen3 models (1.7B-32B), using same-weights Thinking ON/OFF controls; four Hunyuan models provide directional cross-family support. Aggregate pass-rate changes are small (-0.55 to -3.52 pp), yet 10-20% of prompts switch between pass and fail across modes, suggesting that thinking changes the pattern of errors--some prompts improve while others worsen--rather than uniformly degrading performance. Under a post-hoc Qwen3-derived grouping, constraint types separate into Planning (global counting, structure, coordination), which improves at the class level under thinking, and Precision (exact local form), which consistently worsens; the class-level Planning/Precision sign pattern holds directionally for all four Hunyuan models despite Hunyuan's opposite aggregate direction. Thinking also changes final-answer length; matched-length analyses substantially reduce the Precision drop, but a residual penalty remains. Analyzing thinking traces with a cross-encoder relevance metric reveals three patterns: Neutral shows a positive relevance-compliance link (r approximately 0.15); Planning shows near-zero predictive correlation (r approximately 0.02) despite measurable trace engagement, consistent with an execution gap between CE-measured trace relevance and final-answer compliance; Precision shows a small negative correlation (r approximately -0.05), with failing instances having higher mean relevance than passing ones. Activation patching across four model sizes (1.7B-14B) shows that Precision flip instances are more often restored than Planning flip instances (32-58% vs. 14-40% mean layer-restoration), with the largest gap at 14B (about 30 pp).