Current visual representation learning remains bifurcated: vision-language models (e.g., CLIP) excel at global semantic alignment but lack spatial precision, while self-supervised methods (e.g., MAE, DINO) capture intricate local structures yet struggle with high-level semantic context. We argue that these paradigms are fundamentally complementary and can be integrated into a principled multi-task framework, further enhanced by dense spatial supervision. We introduce MTV, a multi-task visual pretraining framework that jointly optimizes a shared backbone across vision-language contrastive, self-supervised, and dense spatial objectives. To mitigate the need for manual annotations, we leverage high-capacity "expert" models--such as Depth Anything V2 and OWLv2--to synthesize dense, structured pseudo-labels at scale. Beyond the framework, we provide a systematic investigation into the mechanics of multi-task visual learning, analyzing: (i) the individual and marginal gain of each objective, (ii) task synergies versus interference, and (iii) scaling behavior across varying data and model scales. Our results demonstrate that MTV achieves "best-of-both-worlds" performance, significantly enhancing fine-grained spatial reasoning without compromising global semantic understanding. Our findings suggest that multi-task learning, fueled by high-quality pseudo-supervision, is a scalable path toward more general visual encoders.
Figures & tables
Figure 1 : Multi-task supervision effects on a ViT-L model with 10M training samples. We integrate VL , SSL , and pseudo-labeled grounding and depth estimation into a unified training framework, which leads to strong and consistent performance gains across diverse tasks.
Figure 2 : Overview of our MTV framework. (a) Each image is paired with a web-crawled caption, and augmented with pseudo region–text pairs and relative depth maps generated by teacher models. (b) MTV jointly learns from three complementary supervision types: Global (image–caption contrast) , Dense (region–text alignment, depth) , and SSL (self-distillation, masked feature prediction) . A shared image encoder is optimized together with a text encoder and an EMA teacher.
Training Tasks
Trainable Parameters
Training Time
VL
203.2 M
1.0×
VL + SSL ⋆
232.6 M (+29.4 M)
1.2× (+20%)
VL + SSL ⋆ + Ground.
239.7 M (+7.1 M)
1.5× (+25%)
VL + SSL ⋆ + Ground. + Depth
250.3 M (+10.6 M)
1.7× (+13%)
Table 1 : Training Cost for a ViT-B/16 model under different training task combinations. (Numbers in parentheses) indicate the increase over the previous row. ⋆ The EMA teacher used in SSL is frozen during training.
Task
Metric
Benchmark
Global Semantic Understanding
Zero-shot Classification
Top-1 Acc.
ImageNet-1k
Zero-shot Retrieval
Recall@1
COCO
Visual Question-Answering
Score
MMVP,
CVBench,
RealWorldQA
Table 2 : Summary of evaluation benchmarks.
ViT
Training Tasks
IN-1k
COCO I
ADE20k
NAVI
SPair
NYUv2 ↓
B-16
VL
55.7
43.0
32.2
39.8
18.0
0.604
VL + SSL
61.4 (+5.7)
48.0 (+5.0)
39.9 (+7.7)
40.4 (+0.6)
22.7 (+4.7)
0.556 (-.048)
VL + Ground.
59.5 (+3.8)
46.8 (+3.8)
37.6 (+5.4)
41.8 (+2.0)
21.2 (+3.2)
0.541 (-.063)
VL + Depth
57.7 (+2.0)
44.0 (+1.0)
36.0 (+3.8)
42.2 (+2.4)
21.3 (+3.3)
0.511 (-.093)
L/16
VL
60.8
46.3
37.3
39.9
19.8
0.575
VL + SSL
65.4 (+4.6)
48.0 (+1.7)
45.5 (+8.2)
43.2 (+3.3)
27.3 (+7.5)
0.501 (-.074)
Table 3 : Individual Task Contributions (50M Data). 10M results are omitted due to space constraints, but show similar trends.
ViT
Data
Training Tasks
IN-1k
COCO
VQA
ADE20k
NAVI
SPair
NYUv2
Acc.
T → I
I → T
Score
mIoU
Recall
Recall
RMSE ↓
B/16
10M
VL
36.2
14.8
21.9
41.1
27.5
39.5
17.0
0.643
VL + SSL
43.7 (+7.5)
19.7 (+4.9)
28.6 (+6.7)
43.1 (+2.0)
36.2 (+8.7)
41.5 (+2.0)
21.1 (+4.1)
0.568 (-.075)
VL + SSL + Ground.
49.0 (+5.3)
23.5 (+3.8)
33.9 (+5.3)
43.2 (+0.1)
39.5 (+3.3)
43.3 (+1.8)
22.4 (+1.3)
0.537 (-.031)
VL + SSL + Ground. + Depth
49.7 (+0.7)
23.9 (+0.4)
34.1 (+0.2)
43.4 (+0.2)
41.7 (+2.2)
43.6 (+0.3)
22.8 (+0.4)
0.512 (-.025)
Absolute Gain Δ
+13.5
+9.1
+12.2
+2.3
+14.2
+4.1
+5.8
-0.131
Table 4 : Marginal Task Gains across data and model scales. (Numbers in parentheses) indicate the incremental increase over the previous row, while Absolute Gain Δ denotes the overall improvement of the full multi-task model over the VL-only baseline.
ViT
Data
Baseline
TaskA
TaskB
Synergy (%)
IN-1k
COCO ⋆
ADE20k
NAVI
SPair
NYUv2
Average
B/16
10M
VL
SSL
Ground.
62.0
64.4
37.9
72.7
31.7
40.8
51.7
SSL
Depth
50.7
41.8
31.0
-10.0
26.8
56.6
32.9
Depth
Ground.
22.8
27.4
33.8
36.4
51.5
56.2
38.1
B/16
50M
VL
SSL
Ground.
26.3
54.0
40.3
100.0
57.4
44.4
53.7
SSL
Depth
8.8
4.0
27.3
75.0
61.7
29.0
34.3
Table 5 : Task Synergy (%) computed as Eq. ( 4 ), measuring the relative gain from combining two tasks beyond the larger of their individual improvements over the VL-only baseline. ⋆ COCO is short for COCO I → T retrieval.
Figure 3 : Data scaling behavior of multi-task visual learning. ViT-Base models are trained with 10M, 50M, and 100M samples under our multi-task setting. CLIP-Base [ 34 ] is shown as gray dashed line. Lower ↓ is better for NYUv2; higher ↑ is better elsewhere.
Figure 4 : Data filtering effects (ViT-B, 10M data × 20 epochs).
ViT
Model
Data
IN-1k
COCO
VQA
ADE20K
NAVI
SPair
NYUv2
Acc.
T → I
I → T
Score
mIoU
Recall
Recall
RMSE ↓
B/16
CLIP [ 34 ]
400M
68.3
33.1
52.4
44.8
42.2
36.9
15.9
0.603
SigLIP [ 50 ]
10B
76.2
47.2
64.5
47.7
45.1
38.5
17.3
0.615
SigLIP2 [ 41 ]
10B
78.2
52.1
68.9
48.1
46.0
38.0
20.8
0.562
MTV
100M
69.5
41.2
57.1
45.6
44.6
43.8
28.4
0.469
L/14
CLIP [ 34 ]
400M
75.5
36.5
56.3
46.5
46.2
36.0
20.5
0.588
Table 6 : Comprehensive comparisons with large-scale VLMs.
School of Computer Science and Technology, Beijing Institute of Technology · 2Zhongguancun Academy · 5Southeast Academy of Information Technology, Beijing Institute of Technology +2