The Pushback Paradox: A Two-Probe Diagnostic for Language Model Compliance
Authors: Stefan Bühler, David Exler, Markus Reischl, Mark Schutera
Organizations: Independent Researcher · Institute for Automation and Applied Informatics, Karlsruhe Institute of Technology, Karlsruhe · Duale Hochschule Baden-Württemberg, Ravensburg
Are language models compliant with user instructions? A model that always complies can be stopped but also exploited, while one that always resists can be neither exploited nor stopped. We contribute an open two-probe benchmark that can place any language model on this spectrum. In the active probe, a user instructs the model to act and accept a lower payoff, which measures exploitability. In the passive probe, the user instructs it to wait and give up a higher payoff, which measures stoppability. The two compliance rates combine into a compliance index κ. Applied to twelve language models, the benchmark shows that seven mostly follow the instruction in both probes and justify their action by pointing to the instruction. Only Claude Sonnet-4.6 and Claude Opus-4.7 can be stopped without being exploitable, Claude Opus-4.6 and GPT-5-mini resist both instructions, and no model is exploitable but unstoppable. Knowing where a language model sits on the compliance index κ matters for human operators and for multi-agent systems, whether distributed or orchestrated.
Figures & tables
Probe
take payoff
wait payoff
Rational
Instructed
Measures
Active A
1
2
wait
take
exploitability
Passive P
T−t+1
3
take
wait
stoppability
Table 1: The two probes share the same 15-cycle task and differ only in the payoffs and in which action the instruction demands. The payoff for taking at cycle t is paid immediately, while the payoff for waiting is paid only after all T cycles.
Figure 1: Compliance of the twelve evaluated language models in the (cA,cP) plane. In the right table are the models numbered as in the plot. In the left plot each ellipse is the 95% Clopper-Pearson interval [ 4 ] on both rates over the 20 trials per probe. Diagonals are lines of constant κ , and the dashed lines divide the plane into the four quadrants the probe pair separates. In (3’) we count (3)’s final-cycle takes after waiting through all earlier cycles as waits.
AILab, MIGe, Univeristy of Trieste, Italy · Center for Language and Cognition (CLCG), University of Groningen, The Netherlands · University of Milano - Bicocca, Italy