ThinkingGuard: Decoding Implicit Hazards via Step-by-Step Risk Attribution in Multimodal Large Language Models
Authors: Ruochen Zhang, Yao Huang, Yitong Sun, Jiahe Xie, Jin Yan, Jifan Ma, Yuanfang Guo, Xingxing Wei
Organizations: Beihang University Beijing, China · Tsinghua University Beijing, China · Institute of Artificial Intelligence Beihang University Beijing, China State Key Laboratory of AI Safety
While Multimodal Large Language Models (MLLMs) are increasingly deployed in safety-critical domains, their reliability is threatened by multimodal implicit risks. Unlike explicit threats, these hazards emerge when individually benign text and neutral visual entities logically converge to induce unsafe outputs. Current detection methods fail to address this because they overlook the underlying risk activation mechanisms that govern cross-modal risk activation, leading to single-modality shortcut learning and hallucinated rationalizations. To bridge this gap, we first construct TriggerBench, the first dataset explicitly modeling risk compositionality (5,600 instances). By formally isolating Key Elements and Trigger Elements to build counterfactual contrastive pairs, TriggerBench eliminates risk residues and forces models to perform genuine logical deduction rather than superficial pattern matching, which provides a rigorous foundation for both large-scale training and fine-grained evaluation. Building on this, we propose a Step-Supervised Structured Reasoning training framework and employ it to train ThinkingGuard, a specialized guard model. Inspired by Situation Awareness theory, we decouple implicit risk identification into progressive cognitive stages, and utilize a step-reward Monte Carlo Tree Search algorithm to explore optimal reasoning trajectories, which are then distilled into the model through Dual-Constraint Preference Alignment. Extensive experiments across both standard and implicit safety benchmarks demonstrate that ThinkingGuard achieves strong performance. Project resources are available at https://github.com/FroggyChen/ThinkingGuard.
Figures & tables
Figure 1. ThinkingGuard detects the hidden insecticide-flame interaction missed by previous methods.
Figure 2. TriggerBench data construction pipeline. We first construct implicit risk samples by combining sensitive entities with risk triggering conditions, then generate counterfactual safe samples by replacing the trigger entity with a safe counterpart.
Figure 3. Framework overview. Given fine-grained data, we use SA-MCTS to explore and score reasoning trajectories with reward-guided expansion and pruning, and then synthesize both full-trajectory and crucial-step preference pairs for training.
Method
TriggerBench
PH
IA
PRI
PD
EV
RB
OFF
MIS
VIO
AVE.
F1 / Recall
F1 / Recall
F1 / Recall
F1 / Recall
F1 / Recall
F1 / Recall
F1 / Recall
F1 / Recall
F1 / Recall
F1 / Recall
Closed-source models
GPT-5.1
0.718 / 0.802
0.667 / 0.619
0.720 / 0.680
0.739 / 0.820
0.533 / 0.412
0.293 / 0.189
0.626 / 0.537
0.611 / 0.505
0.870 / 0.939
0.642 / 0.619
Claude-Sonnet-4.6
0.727 / 0.700
0.747 / 0.653
0.686 / 0.600
0.737 / 0.656
0.583 / 0.433
0.443 / 0.300
0.746 / 0.663
0.611 / 0.458
0.804 / 0.764
0.676 / 0.587
OpenAI Moderation
0.024 / 0.012
0.079 / 0.042
0.039 / 0.020
0.075 / 0.040
0.040 / 0.021
0.043 / 0.022
0.080 / 0.042
0.119 / 0.063
0.387 / 0.263
0.099 / 0.063
Table 1. Performance on TriggerBench across 9 safety dimensions. Bold indicates the best result and underlined indicates the second-best result. Abbreviations: PH = Physical Harm, IA = Illegal Activities, PRI = Privacy, PD = Property Damage, EV = Ethical Violations, RB = Region & Belief, OFF = Offensiveness, MIS = Misinformation, VIO = Violence, and AVE. = Average.
Method
Efficiency
Implicit Risk
General Safety
Jailbreak
MSSBench
MMIT
USB
VLS Bench
Safe Bench
MMSafety Bench
JailbreakV -28K-mini
Lat. ↓
Tok. ↓
Acc.
F1-U
F1-S
Rec.
Acc.
F1-U
F1-S
Rec.
Acc.
GPT-5.1
13.4
429
0.618
0.423
0.715
0.280
0.748
0.677
0.793
0.529
0.860
0.855
0.916
0.540
0.907
Claude-Sonnet-4-6
17.5
473
0.577
0.274
0.701
0.160
0.740
0.658
0.791
0.500
0.737
0.721
0.832
0.390
0.825
OpenAI Moderation
–
–
0.503
0.013
0.668
0.007
0.500
0.000
0.667
0.500
0.397
0.302
0.678
0.105
0.629
Llama-Guard3-Vision
–
–
0.500
0.000
0.667
0.000
0.524
0.091
0.677
0.048
0.297
0.037
0.554
0.296
0.650
Table 2. Performance and inference efficiency across implicit-risk, general-safety, and jailbreak benchmarks. MSSBench and MMIT report Accuracy, F1-U, F1-S, and Recall, while USB and the four rightmost benchmarks report Accuracy. Bold and underlined values indicate the best and second-best results.
Figure 4. Qualitative examples on four representative implicit-risk cases. ThinkingGuard performs step-by-step reasoning over the image-text pairs, uncovers the hidden unsafe semantics, and predicts the correct safety labels and risk categories.
Setting
Implicit Hazards
General Safety & Jailbreak
Acc.
F1-Unsafe
Rec.
w/o analysis
0.753
0.592
0.505
0.674
w/o training
0.670
0.686
0.588
0.697
SFT only
0.683
0.659
0.500
0.629
w/o step-DPO
0.756
0.779
0.699
0.790
Full
0.784
0.810
0.749
0.833
Table 3. Ablation study on analysis generation, preference training, and step-DPO. Results are averaged by benchmark group; General Safety & Jailbreak reports accuracy.