ThinkingGuard: Decoding Implicit Hazards via Step-by-Step Risk Attribution in Multimodal Large Language Models
Authors: Ruochen Zhang, Yao Huang, Yitong Sun, Jiahe Xie, Jin Yan, Jifan Ma, Yuanfang Guo, Xingxing Wei
Organizations: Beihang University Beijing, China · Tsinghua University Beijing, China · Institute of Artificial Intelligence Beihang University Beijing, China State Key Laboratory of AI Safety
While Multimodal Large Language Models (MLLMs) are increasingly deployed in safety-critical domains, their reliability is threatened by multimodal implicit risks. Unlike explicit threats, these hazards emerge when individually benign text and neutral visual entities logically converge to induce unsafe outputs. Current detection methods fail to address this because they overlook the underlying risk activation mechanisms that govern cross-modal risk activation, leading to single-modality shortcut learning and hallucinated rationalizations. To bridge this gap, we first construct TriggerBench, the first dataset explicitly modeling risk compositionality (5,600 instances). By formally isolating Key Elements and Trigger Elements to build counterfactual contrastive pairs, TriggerBench eliminates risk residues and forces models to perform genuine logical deduction rather than superficial pattern matching, which provides a rigorous foundation for both large-scale training and fine-grained evaluation. Building on this, we propose a Step-Supervised Structured Reasoning training framework and employ it to train ThinkingGuard, a specialized guard model. Inspired by Situation Awareness theory, we decouple implicit risk identification into progressive cognitive stages, and utilize a step-reward Monte Carlo Tree Search algorithm to explore optimal reasoning trajectories, which are then distilled into the model through Dual-Constraint Preference Alignment. Extensive experiments across both standard and implicit safety benchmarks demonstrate that ThinkingGuard achieves strong performance. Project resources are available at https://github.com/FroggyChen/ThinkingGuard.
Figures & tables
Figure 1. ThinkingGuard detects the hidden insecticide-flame interaction missed by previous methods.
Figure 2. TriggerBench data construction pipeline. We first construct implicit risk samples by combining sensitive entities with risk triggering conditions, then generate counterfactual safe samples by replacing the trigger entity with a safe counterpart.
Figure 3. Framework overview. Given fine-grained data, we use SA-MCTS to explore and score reasoning trajectories with reward-guided expansion and pruning, and then synthesize both full-trajectory and crucial-step preference pairs for training.
Method
TriggerBench
PH
IA
PRI
PD
EV
RB
OFF
MIS
VIO
AVE.
F1 / Recall
F1 / Recall
F1 / Recall
F1 / Recall
F1 / Recall
F1 / Recall
F1 / Recall
F1 / Recall
F1 / Recall
F1 / Recall
Closed-source models
GPT-5.1
0.718 / 0.802
0.667 / 0.619
0.720 / 0.680
0.739 / 0.820
0.533 / 0.412
0.293 / 0.189
0.626 / 0.537
0.611 / 0.505
0.870 / 0.939
0.642 / 0.619
Claude-Sonnet-4.6
0.727 / 0.700
0.747 / 0.653
0.686 / 0.600
0.737 / 0.656
0.583 / 0.433
0.443 / 0.300
0.746 / 0.663
0.611 / 0.458
0.804 / 0.764
0.676 / 0.587
OpenAI Moderation
0.024 / 0.012
0.079 / 0.042
0.039 / 0.020
0.075 / 0.040
0.040 / 0.021
0.043 / 0.022
0.080 / 0.042
0.119 / 0.063
0.387 / 0.263
0.099 / 0.063
Table 1. Performance on TriggerBench across 9 safety dimensions. Bold indicates the best result and underlined indicates the second-best result. Abbreviations: PH = Physical Harm, IA = Illegal Activities, PRI = Privacy, PD = Property Damage, EV = Ethical Violations, RB = Region & Belief, OFF = Offensiveness, MIS = Misinformation, VIO = Violence, and AVE. = Average.
Method
Efficiency
Implicit Risk
General Safety
Jailbreak
MSSBench
MMIT
USB
VLS Bench
Safe Bench
MMSafety Bench
JailbreakV -28K-mini
Lat. ↓
Tok. ↓
Acc.
F1-U
F1-S
Rec.
Acc.
F1-U
F1-S
Rec.
Acc.
GPT-5.1
13.4
429
0.618
0.423
0.715
0.280
0.748
0.677
0.793
0.529
0.860
0.855
0.916
0.540
0.907
Claude-Sonnet-4-6
17.5
473
0.577
0.274
0.701
0.160
0.740
0.658
0.791
0.500
0.737
0.721
0.832
0.390
0.825
OpenAI Moderation
–
–
0.503
0.013
0.668
0.007
0.500
0.000
0.667
0.500
0.397
0.302
0.678
0.105
0.629
Llama-Guard3-Vision
–
–
0.500
0.000
0.667
0.000
0.524
0.091
0.677
0.048
0.297
0.037
0.554
0.296
0.650
Table 2. Performance and inference efficiency across implicit-risk, general-safety, and jailbreak benchmarks. MSSBench and MMIT report Accuracy, F1-U, F1-S, and Recall, while USB and the four rightmost benchmarks report Accuracy. Bold and underlined values indicate the best and second-best results.
Figure 4. Qualitative examples on four representative implicit-risk cases. ThinkingGuard performs step-by-step reasoning over the image-text pairs, uncovers the hidden unsafe semantics, and predicts the correct safety labels and risk categories.
Setting
Implicit Hazards
General Safety & Jailbreak
Acc.
F1-Unsafe
Rec.
w/o analysis
0.753
0.592
0.505
0.674
w/o training
0.670
0.686
0.588
0.697
SFT only
0.683
0.659
0.500
0.629
w/o step-DPO
0.756
0.779
0.699
0.790
Full
0.784
0.810
0.749
0.833
Table 3. Ablation study on analysis generation, preference training, and step-DPO. Results are averaged by benchmark group; General Safety & Jailbreak reports accuracy.
While Multimodal Large Language Models (MLLMs) show remarkable advancements, their cross-modal capabilities introduce complex vulnerabilities that easily bypass unimodal filters. Existing benchmarks lack fine-grained intent-related annotations and rely on unidimensional metrics, hindering comprehensive robustness evaluation. To address this, we propose MME-Safety, a rigorously verified benchmark featuring a unique four-dimensional annotation schema that categorizes risk scenarios, harm severity, and modality-specific stealth levels. Furthermore, we introduce a hierarchical evaluation framework to assess fundamental response reliability, actual risk exposure, and the structural integrity of defensive behaviors. Extensive zero-shot evaluations across 17 state-of-the-art MLLMs provide a comprehensive safety profile of current multimodal systems. Our analysis systematically investigates cross-modal input configurations and uncovers safety implications associated with Chain-of-Thought (CoT) reasoning. These multifaceted findings underscore the urgent need for robust, reasoning-aware safety alignment in the multimodal landscape.
As large language models (LLMs) are increasingly deployed in real-world applications, safety guardrails are required to go beyond coarse-grained filtering and support fine-grained, interpretable, and adaptable risk assessment. However, existing solutions often rely on rapid classification schemes or post-hoc rules, resulting in limited transparency, inflexible policies, or prohibitive inference costs. To this end, we present YuFeng-XGuard, a reasoning-centric guardrail model family designed to perform multi-dimensional risk perception for LLM interactions. Instead of producing opaque binary judgments, YuFeng-XGuard generates structured risk predictions, including explicit risk categories and configurable confidence scores, accompanied by natural language explanations that expose the underlying reasoning process. This formulation enables safety decisions that are both actionable and interpretable. To balance decision latency and explanatory depth, we adopt a tiered inference paradigm that performs an initial risk decision based on the first decoded token, while preserving ondemand explanatory reasoning when required. In addition, we introduce a dynamic policy mechanism that decouples risk perception from policy enforcement, allowing safety policies to be adjusted without model retraining. Extensive experiments on a diverse set of public safety benchmarks demonstrate that YuFeng-XGuard achieves stateof-the-art performance while maintaining strong efficiency-efficacy trade-offs. We release YuFeng-XGuard as an open model family, including both a full-capacity variant and a lightweight version, to support a wide range of deployment scenarios.
Multimodal Large Language Models are increasingly adopted as autonomous agents in interactive environments, yet their ability to proactively address safety hazards remains insufficient. We introduce SafetyALFRED, built upon the embodied agent benchmark ALFRED, augmented with six categories of real-world kitchen hazards. While existing safety evaluations focus on hazard recognition through disembodied question answering (QA) settings, we evaluate eleven state-of-the-art models from the Qwen, Gemma, and Gemini families on not only hazard recognition, but also active risk mitigation through embodied planning. Our experimental results reveal a significant alignment gap: while models can accurately recognize hazards in QA settings, average mitigation success rates for these hazards are low in comparison. Our findings demonstrate that static evaluations through QA are insufficient for physical safety, thus we advocate for a paradigm shift toward benchmarks that prioritize corrective actions in embodied contexts. We open-source our code and dataset under https://github.com/sled-group/SafetyALFRED.git