cs.CVOct 8, 2026

FearCaut-Qwen: Affective Steering in a Vision-Language Model Shifts the Decision Criterion for Hazard Assessment

Authors: Xiaoshan Zhou

Organizations: School of Project Management, Faculty of Engineering, The University of Sydney, Sydney, NSW 2006, Australia

Abstract

Vision-language models (VLMs) show great potential for damage assessment after a disaster, but a recurring deficiency is that they are reluctant to declare a hazard; that is, recall is low even when overall accuracy appears adequate. This study examines that deficiency by using signal detection theory to decompose the decision behavior into perceptual capability and decision-criterion placement. We then propose a novel method for correcting the over-conservative decision policy, inspired by the finding that fear makes humans risk-averse, and ask whether an affective representation associated with fear can be causally manipulated to similarly alter a VLM's decision tendency. Using mechanistic interpretability, we localize a causally implicated affective circuit in the model and use activation steering to manipulate it while observing the effect on downstream prediction. The method is tested on a two-stage SeisMLLM pipeline built on Qwen2.5-VL-7B-Instruct, which flags only 27.0% of genuinely unsafe buildings on the SeisMLLM-1K test split and never issues a false Red, an SDT criterion of c = +1.354, despite adequate evidence quality (d' = 1.521). An affective direction is localized on emotion-rich natural scenes, causally validated by sparse-neuron knockout and distributed steering on held-out emotion data, and then injected into the building task. Fear-direction injection raises Red recall to 75.7% (p<0.001), and subtracting the same direction suppresses Red predictions entirely, whereas norm-matched random and matched happiness directions show no significant effect. The mechanism is a shift in criterion (c=-1.515) while discrimination is not improved (d'=-0.493). These results show that VLM decisions can be adjusted at inference time without retraining and demonstrate how mechanistic interpretability can be used to diagnose and control VLM decision behaviors in engineering applications.

Explore similar work

CardsList
  1. Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies

    Jul 18, 2026Murali Indukuri, Mohammad Eskandari, Sree Nitya Kollu +2VLM Evaluation

  2. Do VLMs Share Safety Neurons Across Modalities?

    Aug 31, 2026Jiaxuan Li, Jiahao Zhang, Duc Minh Vo +3VLM EvaluationVLM Interpretability

  3. Explore, Map, Remember, Decide: Are Embodied VLMs Ready for Safety-Critical Scenarios?

    Aug 8, 2026Gabriele La Malfa, Nitay Alon, Emanuele La Malfa +2VLM EvaluationMemory-Augmented VLMs