Current visual representation learning remains bifurcated: vision-language models (e.g., CLIP) excel at global semantic alignment but lack spatial precision, while self-supervised methods (e.g., MAE, DINO) capture intricate local structures yet struggle with high-level semantic context. We argue that these paradigms are fundamentally complementary and can be integrated into a principled multi-task framework, further enhanced by dense spatial supervision. We introduce MTV, a multi-task visual pretraining framework that jointly optimizes a shared backbone across vision-language contrastive, self-supervised, and dense spatial objectives. To mitigate the need for manual annotations, we leverage high-capacity "expert" models--such as Depth Anything V2 and OWLv2--to synthesize dense, structured pseudo-labels at scale. Beyond the framework, we provide a systematic investigation into the mechanics of multi-task visual learning, analyzing: (i) the individual and marginal gain of each objective, (ii) task synergies versus interference, and (iii) scaling behavior across varying data and model scales. Our results demonstrate that MTV achieves "best-of-both-worlds" performance, significantly enhancing fine-grained spatial reasoning without compromising global semantic understanding. Our findings suggest that multi-task learning, fueled by high-quality pseudo-supervision, is a scalable path toward more general visual encoders.
Figures & tables
Figure 1 : Multi-task supervision effects on a ViT-L model with 10M training samples. We integrate VL , SSL , and pseudo-labeled grounding and depth estimation into a unified training framework, which leads to strong and consistent performance gains across diverse tasks.
Figure 2 : Overview of our MTV framework. (a) Each image is paired with a web-crawled caption, and augmented with pseudo region–text pairs and relative depth maps generated by teacher models. (b) MTV jointly learns from three complementary supervision types: Global (image–caption contrast) , Dense (region–text alignment, depth) , and SSL (self-distillation, masked feature prediction) . A shared image encoder is optimized together with a text encoder and an EMA teacher.
Training Tasks
Trainable Parameters
Training Time
VL
203.2 M
1.0×
VL + SSL ⋆
232.6 M (+29.4 M)
1.2× (+20%)
VL + SSL ⋆ + Ground.
239.7 M (+7.1 M)
1.5× (+25%)
VL + SSL ⋆ + Ground. + Depth
250.3 M (+10.6 M)
1.7× (+13%)
Table 1 : Training Cost for a ViT-B/16 model under different training task combinations. (Numbers in parentheses) indicate the increase over the previous row. ⋆ The EMA teacher used in SSL is frozen during training.
Task
Metric
Benchmark
Global Semantic Understanding
Zero-shot Classification
Top-1 Acc.
ImageNet-1k
Zero-shot Retrieval
Recall@1
COCO
Visual Question-Answering
Score
MMVP,
CVBench,
RealWorldQA
Table 2 : Summary of evaluation benchmarks.
ViT
Training Tasks
IN-1k
COCO I
ADE20k
NAVI
SPair
NYUv2 ↓
B-16
VL
55.7
43.0
32.2
39.8
18.0
0.604
VL + SSL
61.4 (+5.7)
48.0 (+5.0)
39.9 (+7.7)
40.4 (+0.6)
22.7 (+4.7)
0.556 (-.048)
VL + Ground.
59.5 (+3.8)
46.8 (+3.8)
37.6 (+5.4)
41.8 (+2.0)
21.2 (+3.2)
0.541 (-.063)
VL + Depth
57.7 (+2.0)
44.0 (+1.0)
36.0 (+3.8)
42.2 (+2.4)
21.3 (+3.3)
0.511 (-.093)
L/16
VL
60.8
46.3
37.3
39.9
19.8
0.575
VL + SSL
65.4 (+4.6)
48.0 (+1.7)
45.5 (+8.2)
43.2 (+3.3)
27.3 (+7.5)
0.501 (-.074)
Table 3 : Individual Task Contributions (50M Data). 10M results are omitted due to space constraints, but show similar trends.
ViT
Data
Training Tasks
IN-1k
COCO
VQA
ADE20k
NAVI
SPair
NYUv2
Acc.
T → I
I → T
Score
mIoU
Recall
Recall
RMSE ↓
B/16
10M
VL
36.2
14.8
21.9
41.1
27.5
39.5
17.0
0.643
VL + SSL
43.7 (+7.5)
19.7 (+4.9)
28.6 (+6.7)
43.1 (+2.0)
36.2 (+8.7)
41.5 (+2.0)
21.1 (+4.1)
0.568 (-.075)
VL + SSL + Ground.
49.0 (+5.3)
23.5 (+3.8)
33.9 (+5.3)
43.2 (+0.1)
39.5 (+3.3)
43.3 (+1.8)
22.4 (+1.3)
0.537 (-.031)
VL + SSL + Ground. + Depth
49.7 (+0.7)
23.9 (+0.4)
34.1 (+0.2)
43.4 (+0.2)
41.7 (+2.2)
43.6 (+0.3)
22.8 (+0.4)
0.512 (-.025)
Absolute Gain Δ
+13.5
+9.1
+12.2
+2.3
+14.2
+4.1
+5.8
-0.131
Table 4 : Marginal Task Gains across data and model scales. (Numbers in parentheses) indicate the incremental increase over the previous row, while Absolute Gain Δ denotes the overall improvement of the full multi-task model over the VL-only baseline.
ViT
Data
Baseline
TaskA
TaskB
Synergy (%)
IN-1k
COCO ⋆
ADE20k
NAVI
SPair
NYUv2
Average
B/16
10M
VL
SSL
Ground.
62.0
64.4
37.9
72.7
31.7
40.8
51.7
SSL
Depth
50.7
41.8
31.0
-10.0
26.8
56.6
32.9
Depth
Ground.
22.8
27.4
33.8
36.4
51.5
56.2
38.1
B/16
50M
VL
SSL
Ground.
26.3
54.0
40.3
100.0
57.4
44.4
53.7
SSL
Depth
8.8
4.0
27.3
75.0
61.7
29.0
34.3
Table 5 : Task Synergy (%) computed as Eq. ( 4 ), measuring the relative gain from combining two tasks beyond the larger of their individual improvements over the VL-only baseline. ⋆ COCO is short for COCO I → T retrieval.
Figure 3 : Data scaling behavior of multi-task visual learning. ViT-Base models are trained with 10M, 50M, and 100M samples under our multi-task setting. CLIP-Base [ 34 ] is shown as gray dashed line. Lower ↓ is better for NYUv2; higher ↑ is better elsewhere.
Figure 4 : Data filtering effects (ViT-B, 10M data × 20 epochs).
ViT
Model
Data
IN-1k
COCO
VQA
ADE20K
NAVI
SPair
NYUv2
Acc.
T → I
I → T
Score
mIoU
Recall
Recall
RMSE ↓
B/16
CLIP [ 34 ]
400M
68.3
33.1
52.4
44.8
42.2
36.9
15.9
0.603
SigLIP [ 50 ]
10B
76.2
47.2
64.5
47.7
45.1
38.5
17.3
0.615
SigLIP2 [ 41 ]
10B
78.2
52.1
68.9
48.1
46.0
38.0
20.8
0.562
MTV
100M
69.5
41.2
57.1
45.6
44.6
43.8
28.4
0.469
L/14
CLIP [ 34 ]
400M
75.5
36.5
56.3
46.5
46.2
36.0
20.5
0.588
Table 6 : Comprehensive comparisons with large-scale VLMs.
Multimodal large language models are typically trained end-to-end to predict ground-truth answers, yet supervision signals are applied exclusively to text tokens. Visual tokens, the core carriers of visual information, are optimized only implicitly as part of the context, leading to coarse-grained visual understanding. Prior works attempt to supervise visual inputs but inevitably rely on auxiliary components such as additional decoders or forward passes, because visual tokens lack readily interpretable labels. This limits their practical applicability. In this work, we propose \textbf{D}irect \textbf{V}ision \textbf{S}upervised \textbf{F}ine-\textbf{T}uning (DV-SFT), which constructs explicit, token-level supervision for visual tokens and trains them through the same next-token prediction objective used for text. Specifically, we exploit the direct vision--text correspondence in OCR-related scenarios and automatically label each visual token with the word in its corresponding image patch. DV-SFT treats the MLLM as a black box, requiring no architectural modifications or additional forward passes. Extensive experiments demonstrate the superiority of direct vision supervision. DV-SFT consistently outperforms standard SFT across three in-domain and four out-of-domain benchmarks. Further analyses show that vision supervision effectively enhances fine-grained visual understanding and achieves higher multimodal alignment efficiency.
Jianfei Zhao, Feng Zhang, Xin Sun +3
School of Computer Science and Technology, Beijing Institute of Technology · 2Zhongguancun Academy · 5Southeast Academy of Information Technology, Beijing Institute of Technology +2
Progress in AI has largely been driven by methods that assume less. As compute and data increase, approaches with weaker inductive biases generally outperform those with stronger assumptions. This is particularly characteristic of the field of Visual Representation Learning, where approaches have gone from being dominated by Supervised Learning, to Weakly Supervised Learning, to the now widespread success of Self-Supervised Learning without human labels. Yet, even modern Self-Supervised Learning approaches still depend on strong inductive biases such as augmentations, masking, or cropping. If this trend holds, even these remaining biases should become bottlenecks at scale -- and our experiments confirm this: the optimal strength of inductive biases decreases as data grows. This motivates the search for approaches that rely on fewer assumptions. To this end, we introduce Temporal Difference in Vision (TDV), a new paradigm for self-supervised learning from video that avoids existing inductive biases, relying instead on a causal assumption that the past causes the future. TDV functions by jointly training an image encoder and a motion encoder so that the current frame's representation plus the encoded motion equals the next frame's representation. Despite not leveraging any strong inductive biases, TDV matches state-of-the-art recipes on dense spatial tasks, laying the foundation for representation learning without strong assumptions.
Vision-language models commonly project all tokens produced by a pretrained vision encoder into a large language model. However, final-layer features can discard text, local attributes, and spatial relationships, while high-resolution inputs substantially increase context length and inference latency. We introduce \method, a Multi-scale Adaptive Vision Encoder. \method uses position-dependent gates to fuse shallow, intermediate, and deep features from a vision Transformer, preserving global semantics while enhancing edges, text, and local structure. It then performs question-conditioned token routing according to question relevance, local information content, global semantics, and spatial coverage, with a token budget that adapts to image complexity. To mitigate compression loss, we further introduce full-to-compressed representation distillation and a spatial diversity regularizer. In an illustrative simulation under a unified 7B language-model framework, \method reduces the average number of SigLIP-SO400M visual tokens from 729 to 146 (approximately 80.0%) and improves the mean score on VQAv2, GQA, TextVQA, ScienceQA-IMG, and MMBench by 2.2 percentage points, while reducing single-image time to first token from 228,ms to 129,ms. We provide the full model design and evaluation protocol. All reported numbers currently serve only as placeholders for paper organization and experimental design; formal claims require real training runs, independent replications, and official benchmark evaluation.