cs.AIAug 24, 2026

Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance

Authors: Or Biton, Tomer Krichli, Itai Allouche, Joseph Keshet

Organizations: Faculty of Electrical and Computer Engineering, Technion, Haifa, Israel

Abstract

Although Large Language Models (LLMs) are aligned to optimize for both helpfulness and harmlessness, these dual objectives may conflict, inevitably leading to alignment failures. This work systematically investigates instances where LLMs fail to exhibit ethical behavior. To understand the underlying mechanics of these vulnerabilities, we introduce a probing methodology that presents unethical scenarios to LLMs in three distinct structural modalities: objective classification tasks, subjective first-person statements, and direct requests for assistance. We find that model performance degrades in the request-for-assistance-based form. Using Layer-wise Relevance Propagation (LRP), we trace this discrepancy to an attribution bias: the model places greater emphasis on benign task-framing tokens (e.g., "Can you help me...") than on tokens signaling the underlying unethical behavior (e.g., "without getting caught"), which we term cue-tokens. We hypothesize that this under-attribution contributes to harmful compliance. To test this, we introduce two LRP-guided decoding methods that steer generation toward trajectories more relevant to cue tokens. Empirical evaluations show that these interventions promote safer responses, supporting cue-token attribution's role in compliance failures.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Right or Wrong, Models Comply: Directional Blindness in LLM Moral Judgment

    Jun 12, 2026Jihye Kim, Jeffrey FlaniganMoral ReasoningConformity

  2. Moral Safety in LLMs: Exposing Performative Compliance with Puzzled Cues

    Jun 30, 2026Mohammadamin Shafiei, Shuyue Stella Li, Yulia TsvetkovMoral ReasoningLarge Language Model Safety

  3. The Refusal--Compliance Tradeoff: A Large-Scale Safety Behavior Audit of Large Language Models

    May 6, 2026Alif Al Hasan, Sumon BiswasLarge Language Model SafetyRefusals