OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Authors: Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, +8 more
Organizations: University of California, Berkeley · Duke University · Impossible Research · Carnegie Mellon University · University of Washington · Elorian
Training a model to generate visual content can encourage it to learn rich perceptual capabilities related to geometry, spatial relationships, and objectness; yet, its benefits for visual understanding remain unclear. We ask: when and how does visual generation supervision improve visual understanding? We study controlled pairs of image-to-image (I2I) generation and image-to-text (I2T) understanding tasks that express the same underlying problem in different output modalities. We find that under the correct recipe, I2I training improves downstream I2T performance, with larger gains as the amount of I2I training data increases. We next ask which generation tasks benefit which understanding capabilities. To study transfer beyond paired tasks, we introduce OmniTaskonomy, a unified taxonomy spanning 19 I2I generation tasks and 25 I2T understanding capabilities. The resulting transfer map reveals selective, task-dependent benefits. Some follow intuitive correspondences, e.g., depth prediction improving metric 3D reasoning, object pointing improving counting, and jigsaw reconstruction improving 2D ordering. Interestingly, we also uncover surprising connections: 2.5D segmentation improving category recognition and Z-depth prediction improving localization. To probe these patterns, we analyze gradient alignment between generation and understanding tasks and find that stronger alignment is associated with larger downstream transfer gains. Together, our results highlight visual generation as a rich source of supervision for visual understanding and provide a roadmap for unlocking its benefits through the right training curriculum and task selection. Project page: https://omni-taskonomy.github.io/.
Figures & tables
Figure 1: When and how does visual generation improve visual understanding? (a) I2I training followed by I2T fine-tuning yields increasing gains as more I2I data is added, whereas mixed training shows no consistent scaling benefit. (b) Different generation tasks (I2I) benefit different understanding capabilities. Blue and red indicate accuracy gains and drops relative to I2T-only training (percentage points). (c) Higher average gradient alignment in the understanding branch’s pre-attention RMSNorm parameters is associated with larger mean transfer gains across capabilities.
Figure 2: Scaling I2I and I2T supervision. (a) Controlled Jigsaw and Zoom-In tasks with image and text outputs. (b) I2T accuracy versus I2I training examples across recipes, with 1k I2T examples. (c) I2T accuracy versus I2T training examples for I2T-only and I2I → I2T training (100k I2I examples). Curves show means over three seeds. Shading indicates ±1 SEM.
Jigsaw
Zoom-In
I2I Acc.
70.9
91.8
I2T Acc.
81.3
90.6
P(I2T✓∣I2I✓)
88.2
90.9
P(I2T✓∣I2I×)
64.3
86.7
Table 1: I2I–I2T instance alignment.
Figure 3: OmniTaskonomy. A unified taxonomy of I2I tasks and understanding capabilities under the three Rs: Recognition , Reconstruction , and Reorganization . I2I tasks and understanding capabilities occupy separate leaves in a shared hierarchy.
Figure 4: Generation-to-understanding transfer across visual capabilities. Each column corresponds to an I2I source task and each row to an understanding capability in OmniTaskonomy. Cells report the change in I2T accuracy, in percentage points, relative to the I2T-only baseline. Results are averaged over three seeds. Blue indicates positive transfer and red indicates negative transfer. Cells with p<0.05 under an exact paired permutation test against the I2T-only baseline are outlined and labeled with their gains. The detailed numbers in the heat map are in Appendix D.2 .
Figure 5: Gradient alignment and downstream transfer. (a,b) Gradient alignment for the controlled Jigsaw and Zoom-In pairs across model components and individual pre-attention RMSNorm layers. (c) Mean gradient alignment versus mean transfer for seven understanding capabilities, averaging over 19 I2I sources. (d) Alignment and transfer for all 133 source–target pairs.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Optimizer
AdamW with β1=0.9 , β2=0.95 , ϵ=10−15
Peak learning rate
2×10−5
Weight decay
0
Schedule
Linear warm-up, then constant learning rate
Warm-up
50 updates for initial I2I training; 8 for I2T or Mixed training
Gradient clipping
Global gradient ℓ2 norm capped at 1.0
Appendix
Table 2: Optimization settings.
Plot label
I2I examples
Mixed batch (I2I:I2T)
3k
1,876
4:64
10k
9,380
20:64
30k
30,016
64:64
100k
100,000
212:64
Appendix
Table 3: I2I pool sizes and mixed batches.
Figure 6: BAGEL architecture.
Figure 7: OmniTaskonomy construction pipeline. The three stages are schema annotation, taxonomy refinement, and VLM verification.
Figure 8: From an example to a shared schema. Free-form and fixed-schema attributes come from separate calls on the same example.
Stage
Recorded assignment
Local tree
Natural and Surveillance Photos → Counting and Enumeration → Object Counting → Small Count Enumeration
Global candidate tree
Natural Images → Quantitative and Logical Reasoning → Object Counting → Small Count Enumeration
Final 3R label
Reorganization → Counting
Appendix
Table 4: A counting example through taxonomy construction.
Capability
Benchmark question
Routing rationale
Appearance understanding
MMStar: “What is the dominant color in the image?”
Queries a named appearance property.
Depth understanding
BLINK: “Which point is closer to the camera?”
Compares camera-relative depth.
Metric 3D relation
CV-Bench-3D: “Which highlighted object is closer to the car in real-world distance?”
Requires metric inter-object proximity.
2D spatial relation
CV-Bench-2D: “Where is the mouse relative to the potted plant?”
Queries their visible left/right arrangement.
Connectivity
BLINK: “Is the bed touching the book?”
Queries contact rather than metric distance.
Counting
CV-Bench-2D: “How many trees are in the image?”
Requires cardinality of a visually selected set.
Appendix
Table 5: Representative retained annotations. Questions are shortened for space; rationales summarize the recorded assignments.
Reviewers
n
Agree (%)
κ
R1–R2
48
97.92
0.98
R1–R3
47
95.74
0.95
R1–R4
47
97.87
0.98
R2–R3
48
91.67
0.91
R2–R4
48
93.75
0.93
R3–R4
47
95.74
0.95
Appendix
Table 6: Pairwise agreement between human reviewers. Each comparison excludes examples that either reviewer marked unsure ; n is the number of remaining shared examples.
Capability
N
Description
Recognition
Category recognition
962
Recognize the semantic category, name, stable identity attribute, or known identity of an object, person, place, or scene.
Part recognition
54
Identify or name a sub-part / component of an object (a handle, a wing, an engine bay).
Appearance understanding
687
Understand visible appearance in the image plane, including named non-geometric properties and projected 2D shape, silhouette, outline, contour morphology, or edge structure.
Visual similarity
248
Judge holistic visual resemblance, difference, or match among complete images, regions, or objects when no localized indexed correspondence and no specific appearance property is requested.
State recognition
340
Recognize an ordinary temporary, operational, environmental, or emotional state without judging it against a norm.
Appendix
Table 7: Understanding capabilities in OmniTaskonomy. N is the number of retained evaluation examples for each capability.
I2I task
Target image
Recognition
Object editing
An image with objects of a requested category added or replaced.
Attribute editing
An image with a specified color, material, or other appearance attribute changed.
Reconstruction
Colorization
The color image corresponding to a grayscale input.
Z-depth
Depth along the camera’s viewing axis.
Appendix
Table 8: I2I tasks and their supervised targets.
Benchmark
Collected
Retained
Excluded
Retained (%)
BLINK
1901
1758
143
92.5
CV-Bench-2D
1438
1434
4
99.7
CV-Bench-3D
1200
1200
0
100.0
MMStar
1500
986
514
65.7
MMT-Bench
3127
2840
287
90.8
MMVP
300
290
10
96.7
Appendix
Table 9: Benchmark coverage. CV-Bench is reported as separate 2D and 3D evaluation sets.
Figure 9: Capability coverage and benchmark composition. Bar lengths give the number of retained evaluation examples per capability, with totals labeled, and colors indicate the source benchmark. The 19 capabilities with more than 100 examples appear in the main transfer map (Figure 4 ); Figure 10 includes all 25.
Task
Source
Recognition
Object editing
ConceptEdit-12M, category-routed pairs
Attribute editing
ConceptEdit-12M, appearance-routed pairs
Reconstruction
Colorization
Taskonomy tiny
Z-depth
Taskonomy tiny
Appendix
Table 10: I2I training sources. Each task uses 50,000 I2I training examples.
Setting
Stage 1 (I2I)
Stage 2 (I2T)
Training-example budget
50,000
50,000
Batch composition
I2I only
I2T only
GPUs / effective batch size
4 / 64
4 / 64
Optimizer updates
781–782
781
Optimizer
AdamW, β=(0.9,0.95)
Learning rate
2×10−5
Appendix
Table 11: Transfer training settings.
Capability
Base (%)
Capability
Base (%)
Category recognition
75.09
Metric 3D
68.65
Part recognition
74.07
Occlusion
78.12
Appearance
66.13
3D shape reasoning
31.11
Visual similarity
77.96
Orientation
55.72
State recognition
68.73
2D spatial
88.21
Activity
72.20
Multi-view
50.56
Appendix
Table 12: I2T-only baseline accuracy on each understanding capability, averaged over three random seeds.
Figure 10: Complete transfer matrix. Panel (a) shows the Recognition (2) and Reconstruction (9) sources, and panel (b) the Reorganization (8) sources, including all 25 target capabilities. Values are accuracy changes over the I2T-only baseline in percentage points. Outlined borders mark exact p<0.05 from paired permutation tests (Appendix D.3 ). Parentheses give the number of evaluation examples.
Understanding capability
Evaluation examples
Mean alignment
Mean gain (pp)
Category recognition
962
0.108
+0.44
Appearance understanding
687
0.019
+0.01
Depth understanding
820
−0.038
+0.07
Metric 3D relation
723
−0.005
+1.75
2D spatial relation
927
−0.001
−0.40
Counting
1,542
0.186
+1.32
Appendix
Table 14: Understanding capabilities in the cross-task gradient analysis. Each capability has more than 500 evaluation examples, of which 500 are used for gradient sampling. Alignment and transfer gain are averaged over the 19 I2I sources, as in Figure 5 (c).
Unified Multimodal Models (UMMs) integrate image understanding and generation within a single architecture, yet how the two tasks interact remains understudied. We investigate transferability in UMMs: whether training a capability on one task improves the same capability on the other without explicit supervision. Through controlled experiments, we empirically find that transferability depends on architecture-models with fully shared transformer backbone and a unified visual encoder exhibit consistent cross-task transfer, while loosely coupled designs show little or none. Leveraging this transferability, we propose a practical training strategy. The most straightforward way to improve a target generative capability (e.g., counting) is to fine-tune generation directly, but this can degrade visual quality due to distribution shift. Instead, we train the corresponding understanding task and let it transfer into generation, which improves capability-specific generative performance while minimizing distribution shift. We validate this across three capabilities-counting, spatial relation, and text recognition/generation-showing that cross-task transferability can be systematically exploited in UMMs.
Jiwon Kang, Heeji Yoon, Jaewoo Jung +5
KAIST AI, South Korea · Blynx, South Korea · Trillionlabs, South Korea
Recent works show that image and video generators exhibit zero-shot visual understanding behaviors, in a way reminiscent of how LLMs develop emergent capabilities of language understanding and reasoning from generative pretraining. While it has long been conjectured that the ability to create visual content implies an ability to understand it, there has been limited evidence that generative vision models have developed strong understanding capabilities. In this work, we demonstrate that image generation training serves a role similar to LLM pretraining, and lets models learn powerful and general visual representations that enable SOTA performance on various vision tasks. We introduce Vision Banana, a generalist model built by instruction-tuning Nano Banana Pro (NBP) on a mixture of its original training data alongside a small amount of vision task data. By parameterizing the output space of vision tasks as RGB images, we seamlessly reframe perception as image generation. Our generalist model, Vision Banana, achieves SOTA results on a variety of vision tasks involving both 2D and 3D understanding, beating or rivaling zero-shot domain-specialists, including Segment Anything Model 3 on segmentation tasks, and the Depth Anything series on metric depth estimation. We show that these results can be achieved with lightweight instruction-tuning without sacrificing the base model's image generation capabilities. The superior results suggest that image generation pretraining is a generalist vision learner. It also shows that image generation serves as a unified and universal interface for vision tasks, similar to text generation's role in language understanding and reasoning. We could be witnessing a major paradigm shift for computer vision, where generative vision pretraining takes a central role in building Foundational Vision Models for both generation and understanding.
While Multimodal Large Language Models (MLLMs) have achieved remarkable progress, visual understanding and generation are typically treated as divergent objectives. Existing unified frameworks often rely on discrete visual tokenization or diffusion objectives whose generative targets differ from the continuous representations consumed by visual understanding models, making direct transfer to enhance existing pretrained MLLMs non-trivial. In this work, we present GAS, a generation-guided training framework that reinterprets visual generation as auxiliary supervision for representation learning. Concretely, GAS adapts Next Embedding Prediction (NEP) as a cross-modal generation paradigm within a decoupled Mixture-of-Transformers (MoT) architecture. By maintaining a shared lower trunk and parallel upper layers, GAS lets generation losses enrich the shared visual pathway with finer spatial precision and stronger visual retention while shielding the upper understanding layers from direct generation gradients. To maximize this synergy, we further construct highly correlated generation tasks that demand deep cognitive grounding rather than generic synthesis alone. Across model scales and training stages, GAS improves aggregate multimodal understanding, with its most reliable gains on perception and spatial comprehension. Crucially, because the auxiliary generation branch is discarded after training, these gains incur zero inference overhead. Extensive controlled comparisons and representation-level analyses further clarify when and why generation-guided training benefits understanding, and demonstrate the feasibility of generation-guided training as a practical route to stronger multimodal understanding.