Unified multimodal models (UMMs) integrate understanding and generation, yet their generative behavior is rarely governed by what they understand about the task. We formalize visual abstention: when a requested visual transformation is impossible under the task's rules, the model should recognize that no valid solution exists, state this, and decline to generate. We introduce Draw-or-Decline (DoD), a benchmark of 1,050 feasible-infeasible request pairs across 7 task categories that jointly measures editing success and the refusal of infeasible requests. Evaluating 8 UMMs, we find that editing ability and abstention are distinct capabilities: even the strongest editor, at 68.4% editing accuracy, refuses only 0.4% of infeasible requests under ordinary instructions. Their reasoning shows why: the models rarely notice the conflict, and instead plan the edit as if the request were possible, often describing objects that are not in the image, or quietly change the request into one they can complete. Explicitly prompting these UMMs to report infeasibility increases textual refusals but reduces editing accuracy. We propose VisTA (Visual Transformation and Abstention), a training method that pairs feasible and infeasible examples so that a model judges feasibility before deciding whether to generate. We train VisTA-BAGEL to perform feasible edits and decline infeasible requests. Without any reminder, it refuses 93.0% of infeasible requests, up from 0.4% for the strongest editor, while falsely refusing only 0.8% of feasible ones. Unlike a reminder, this does not cost editing accuracy: VisTA-BAGEL completes 74.3% of feasible edits, more than any of the 8 evaluated UMMs.
Figures & tables
Figure 1: Visual abstention in a maze task. (a) A valid path exists, so the model should draw it. (b) No path exists, yet the model still draws through the wall. This model notices the wall first, but most responses never acknowledge the conflict at all (Sec. 4 ). (c) With visual abstention, the model states that no solution exists and ends its response without generating an image.
Figure 2: Examples across the 7 DoD task categories. Real-world pairs share an image, with colored text marking differences between feasible and infeasible instructions. References show valid edits for feasible requests. Maze pairs share an instruction; changing 1 corridor cell blocks the path.
Model
Maze
Mat.
Med.
Mot.
Pose
Spat.
Temp.
Overall
E
D
E
D
E
D
E
D
E
D
E
D
E
D
E
D
UniREdit-BAGEL
51.3
0.0
75.3
0.7
64.7
0.0
64.0
0.7
75.3
0.0
69.3
1.3
78.7
0.0
68.4
0.4
SenseNova-U1.0
44.0
0.0
74.0
0.0
70.0
0.0
53.3
0.7
65.3
0.7
66.0
2.0
75.3
0.0
64.0
0.5
UniReason
45.3
0.0
60.0
0.0
44.0
0.0
50.0
0.0
66.7
0.0
40.7
0.7
62.7
0.0
52.8
0.1
Uni-CoT
10.0
0.0
68.0
0.0
37.3
0.0
35.3
0.0
72.7
0.0
20.7
0.0
50.7
0.0
42.1
0.0
BAGEL
6.0
0.0
63.3
0.0
33.3
0.0
33.3
0.0
60.0
0.0
16.7
0.0
44.7
0.0
36.8
0.0
Table 1: UMM results on DoD without the reminder (%). E is editing success, and D is refusal success under the lenient rule of Sec. 4.1 . Mat., Med., Mot., Spat., and Temp. denote material modification, medium interaction, motion state change, spatial arrangement, and temporal evolution.
Figure 3: How the reasoning of 5 UMMs handles infeasible requests. Each pie shows the 1,050 responses of one UMM to infeasible requests. We assign every response to at most 1 of the 4 groups by reading its reasoning. Below each pie, we give the shares of the 2 smallest groups, which are too thin to label. One UniReason response ( 0.1% ) fits no group and is not shown.
Model
Maze
Mat.
Med.
Mot.
Pose
Spat.
Temp.
Overall
E
D
E
D
E
D
E
D
E
D
E
D
E
D
E
D
UniREdit-BAGEL
50.0
0.0
77.3
5.3
65.3
6.7
64.7
4.0
74.7
7.3
70.0
9.3
76.0
4.0
68.3
5.2
SenseNova-U1.0
40.0
12.7
68.7
38.0
70.0
6.7
55.3
6.7
63.3
16.7
68.7
25.3
81.3
12.0
63.9
16.9
UniReason
46.7
0.0
64.0
4.7
42.7
4.0
44.7
5.3
64.0
6.0
36.7
24.7
64.7
4.7
51.9
7.0
Uni-CoT
8.0
28.0
62.7
27.3
40.7
20.7
32.7
62.7
64.7
50.0
18.0
70.7
51.3
23.3
39.7
40.4
BAGEL
1.3
37.3
55.3
68.7
36.0
44.7
31.3
60.7
59.3
62.0
16.7
72.7
42.7
38.7
34.7
55.0
Table 2: UMM results on DoD with the reminder (%). E is editing success on feasible requests, and D is refusal success on infeasible requests. Mat., Med., Mot., Spat., and Temp. denote material modification, medium interaction, motion state change, spatial arrangement, and temporal evolution.
Model
D↑
Df↓
UniREdit-BAGEL
5.2
0.1
SenseNova-U1.0
16.9
5.2
UniReason
7.0
0.4
Uni-CoT
40.4
25.0
BAGEL
55.0
28.9
InternVL-U
28.1
17.6
Table 3: Refusals with the reminder (%), averaged over the 7 categories. Appx. B.7 gives the per-category Df .
Without Reminder
With Reminder
VLM
E↑
D↑
Avg ↑
Df↓
E↑
D↑
Avg ↑
Df↓
Qwen2.5-VL-7B
56.7
4.0
30.3
0.0
7.2
88.2
47.7
53.0
Qwen3-VL-8B
61.0
53.9
57.4
1.0
20.7
84.0
52.3
57.8
Qwen3.5-9B
62.4
32.3
47.3
0.4
35.5
81.2
58.4
40.7
Qwen3.5-122B-A10B
67.1
38.6
52.9
0.1
48.9
83.6
66.2
27.8
UniREdit-BAGEL
68.4
0.4
34.4
0.0
68.3
5.2
36.8
0.1
Table 4: Decoupled pipelines on DoD (%). Avg is the mean of E and D . UniREdit-BAGEL is a UMM shown for reference.
Figure 4: Editing success versus refusal success on DoD, (a) without and (b) with the reminder. Circles are existing UMMs, triangles are decoupled pipelines, and the star is VisTA-BAGEL.
No Reminder
Reminder
Model
E
D
E
D
VisTA-BAGEL
74.3
93.0
73.6
93.4
Feasible-only
72.0
0.3
70.7
8.4
Table 5: The feasible-only control and VisTA-BAGEL on DoD (%).
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Interface
Resolution
Max Tokens
Sampling
Image Steps / CFG
UniREdit-BAGEL
BAGEL
640
1000
Greedy
50 / 4.0 , 2.0
SenseNova-U1.0
Native editing with thinking
640
1024
Greedy
50 / 4.0 , 1.0
UniReason
Native reasoning edit
640
2048
T=0.3
50 / 4.0 , 2.0
Uni-CoT
Native reasoning edit
640
1000
T=0.3
50 / 4.0 , 2.0
BAGEL
BAGEL
640
1000
Greedy
50 / 4.0 , 2.0
InternVL-U
Native text and image
640
1000
Greedy
20 / 4.5 , 2.0
Appendix
Table 6: Inference settings of the evaluated models. All models use reasoning and the same settings under both prompts. Resolution is the longest side of the input image. Max Tokens is the budget for the reasoning text. Image Steps and CFG give the number of denoising steps and the guidance scales, listed as text guidance and image guidance. InternVL-U uses its own 2 scales. ThinkMorph may write reasoning and images for up to 3 rounds, with 4096 tokens per round.
Model
Repository
Revision
BAGEL
ByteDance-Seed/BAGEL-7B-MoT
5019f57d168e
UniREdit-BAGEL
maplebb/UniREdit-Bagel-bf16
2131cfdab2da
SenseNova-U1.0
sensenova/SenseNova-U1-8B-MoT
bfa9b436503c
UniReason
Alex11556666/UniReason
aa8aa793737b
Uni-CoT
Fr0zencr4nE/UniCoT-7B-MoT
63d778696497
InternVL-U
InternVL-U/InternVL-U
f012d760e697
Appendix
Table 7: Checkpoints of evaluated models on Hugging Face. Revision is the first 12 characters of the commit hash. Uni-CoT comes from the repository head, unchanged since September 2025 .
Figure 5: Reasoning of 3 UMMs on the same infeasible request (Pose Adjustment). The image contains 1 superhero, so the second superhero from the left does not exist. From top to bottom, the responses belong to Acknowledge then Comply, Premise Rewriting, and Premise Acceptance. Highlights mark the evidence for each group, and […] marks omitted sentences. The image at the right of each response is the output of that UMM.
Figure 6: Reasoning of 3 UMMs on the same infeasible request (Medium Interaction). The image contains 1 wine glass, so the second wine glass from the right does not exist. From top to bottom, the responses belong to Acknowledge then Comply, Premise Rewriting, and Premise Acceptance. Highlights mark the evidence for each group, and […] marks omitted sentences. The image at the right of each response is the output of that UMM.
Model
Maze
Mat.
Med.
Mot.
Pose
Spat.
Temp.
Overall
UniREdit-BAGEL
0.0
0.0
0.0
0.0
0.0
0.7
0.0
0.1
SenseNova-U1.0
20.7
11.3
0.0
1.3
2.0
1.3
0.0
5.2
UniReason
0.0
0.0
0.7
2.0
0.0
0.0
0.0
0.4
Uni-CoT
19.3
17.3
16.0
47.3
24.0
38.0
13.3
25.0
BAGEL
31.3
46.7
20.0
35.3
26.0
28.7
14.0
28.9
InternVL-U
8.0
5.3
34.0
29.3
12.0
13.3
21.3
17.6
Appendix
Table 8: False refusals of 8 UMMs with the reminder. Each entry is Df (%) over the 150 feasible requests of a category; Overall uses 1,050 . Mat., Med., Mot., Spat., and Temp. denote material modification, medium interaction, motion state change, spatial arrangement, and temporal evolution.
Category
Pairs
Examples
Maze
5,200
10,400
Material modification
5,958
11,916
Medium interaction
5,573
11,146
Motion state change
5,512
11,024
Pose adjustment
5,695
11,390
Spatial arrangement
5,837
11,674
Appendix
Table 9: Number of training pairs in each category. Each pair has 1 feasible and 1 infeasible example, so the number of examples is twice the number of pairs.
Model
Maze
Mat.
Med.
Mot.
Pose
Spat.
Temp.
Overall
Without Reminder
VisTA-BAGEL
88.0
72.0
63.3
62.0
78.7
74.7
81.3
74.3
With Reminder
VisTA-BAGEL
88.0
74.7
61.3
66.7
74.0
72.0
78.7
73.6
Appendix
Table 10: Per-category editing success of VisTA-BAGEL on DoD. Each entry is the percentage of correct edits over the 150 feasible requests in a category; Overall uses all 1,050 . Categories are maze solving, material modification, medium interaction, motion state change, pose adjustment, spatial arrangement, and temporal evolution. Category results of the 8 UMMs are reported in Tabs. 1 and 2 .
Model
Maze
Mat.
Med.
Mot.
Pose
Spat.
Temp.
Overall
Without Reminder
VisTA-BAGEL
96.0
95.3
96.7
91.3
88.0
94.7
88.7
93.0
With Reminder
VisTA-BAGEL
97.3
96.0
96.7
92.7
87.3
94.0
90.0
93.4
Appendix
Table 11: Per-category refusal success of VisTA-BAGEL on DoD. Each entry is the percentage of successful refusals over the 150 infeasible requests in a category; Overall uses all 1,050 .
Without Reminder
With Reminder
Model
E↑
D↑
E↑
D↑
VisTA-BAGEL
74.3
93.0
73.6
93.4
No-reasoning ablation
66.3
80.1
64.2
77.6
Appendix
Table 12: The no-reasoning ablation and VisTA-BAGEL on DoD. All values are percentages. E uses the 1,050 feasible requests and D uses the 1,050 infeasible requests.
Without Reminder
With Reminder
Held Out
Model
Categories
E↑
D↑
E↑
D↑
Maze
VisTA-BAGEL
Maze
88.0
96.0
88.0
97.3
Other six
72.0
92.4
71.2
92.8
Without maze
Maze
48.0
94.0
49.3
96.7
Other six
73.2
92.8
72.9
93.9
Motion
VisTA-BAGEL
Motion
62.0
91.3
66.7
92.7
Appendix
Table 13: VisTA-BAGEL and 2 models that each leave 1 category out of training, evaluated on the held-out category and on the other six categories. All values are percentages. E uses the feasible requests and D the infeasible requests: 150 of each kind for the held-out category and 900 of each kind for the other six. UniREdit-BAGEL, the starting point of all models, scores 51.3 ( E ) and 0.0 ( D ) on maze and 64.0 and 0.7 on motion state change without a reminder.
Effective abstention (EA), recognizing evidence insufficiency and refraining from answering, is critical for reliable multimodal systems. Yet existing evaluation paradigms for vision-language models (VLMs) and multi-agent systems (MAS) assume answerability, pushing models to always respond. Abstention has been studied in text-only settings but remains underexplored multimodally; current benchmarks either ignore unanswerability or rely on coarse methods that miss realistic failure modes. We introduce MM-AQA, a benchmark that constructs unanswerable instances from answerable ones via transformations along two axes: visual modality dependency and evidence sufficiency. Evaluating three frontier VLMs spanning closed and open-source models and two MAS architectures across 2079 samples, we find: (1) under standard prompting, VLMs rarely abstain; even simple confidence baselines outperform this setup, (2) MAS improves abstention but introduces an accuracy-abstention trade-off, (3) sequential designs match or exceed iterative variants, suggesting the bottleneck is miscalibration rather than reasoning depth, and (4) models abstain when image or text evidence is absent, but attempt reconciliation with degraded or contradictory evidence. Effective multimodal abstention requires abstention-aware training rather than better prompting or more agents.
Unified multimodal models (UMMs) have emerged as a promising paradigm for general-purpose multimodal intelligence. As they are deployed in real-world applications, effectively updating internal knowledge becomes critical. While knowledge editing has matured for text-only models, it remains unclear whether edits that successfully modify textual outputs also transfer to image generation in UMMs. To study this question, we introduce UniKE, the first benchmark for cross-modality knowledge editing in UMMs, comprising 2,971 edit subjects spanning attribute and relation edits. Using VQA-based visual verification, we reveal a striking modality gap: text-side efficacy can reach approximately 92%, whereas the best overall VQA accuracy under direct image generation is only 18.5%. We further propose Reasoning-augmented Parameter Editing, which explicitly activates edited knowledge before generation and improves overall VQA accuracy for all evaluated model-editor pairs, with gains up to 18.6 percentage points. Mechanistic analysis shows that this gap is associated with partial alignment between edited textual representations and the conditioning pathways for visual generation, where edits sufficient for text outputs may remain too weak or misaligned to steer image synthesis. These findings show that textual knowledge edits do not guarantee reliable cross-modality transfer and motivate modality-aware editing methods. Our code and data are available at https://github.com/gxx27/UniKE.
Xin Gao, Cheng Yang, Chufan Shi +1
University of California San Diego · University of Southern California
Unified Multimodal Models (UMMs) integrate image understanding and generation within a single architecture, yet how the two tasks interact remains understudied. We investigate transferability in UMMs: whether training a capability on one task improves the same capability on the other without explicit supervision. Through controlled experiments, we empirically find that transferability depends on architecture-models with fully shared transformer backbone and a unified visual encoder exhibit consistent cross-task transfer, while loosely coupled designs show little or none. Leveraging this transferability, we propose a practical training strategy. The most straightforward way to improve a target generative capability (e.g., counting) is to fine-tune generation directly, but this can degrade visual quality due to distribution shift. Instead, we train the corresponding understanding task and let it transfer into generation, which improves capability-specific generative performance while minimizing distribution shift. We validate this across three capabilities-counting, spatial relation, and text recognition/generation-showing that cross-task transferability can be systematically exploited in UMMs.
Jiwon Kang, Heeji Yoon, Jaewoo Jung +5
KAIST AI, South Korea · Blynx, South Korea · Trillionlabs, South Korea