Organizations: Department of Intermedia Art and Science, Waseda University, Tokyo, Japan · SB Intuitions Corp., Tokyo, Japan · Department of Electronic and Physical Systems, Waseda University, Tokyo, Japan · Department of Computer Science and Engineering, Waseda University, Tokyo, Japan · National Institute of Advanced Industrial Science and Technology (AIST), Tokyo, Japan
Language-conditioned robot policies have made clear progress in multitask manipulation, but task-relevant local visual evidence usually stays hidden inside a visual backbone or attention layers. This leaves the policy difficult to inspect and fragile under visual change, two symptoms of a missing explicit, task-aligned local visual channel. We present TLC-DiT, a plug-in extension of the Multitask Diffusion Transformer (DiT) policy that adds explicit task-guided local visual feature maps without changing the diffusion objective or the action-generation process. For each camera view, frozen DINOv2 patch features are modulated by the CLIP task embedding through FiLM and refined by a lightweight CoordConv CNN adapter into smooth spatial maps, which are concatenated with the original global image, language, joint-state, and timestep conditions. On LIBERO, TLC-DiT reaches a 93.5% average success rate, compared with 86.5% for Multitask DiT and 79.25% for SmolVLA. On LIBERO-plus, the total success rate improves from 54.07% to 57.24%, with larger gains under camera, background, and sensor-noise changes. In real-world bimanual tasks, TLC-DiT raises Teabag Putting completion from 44% to 89% while maintaining comparable Match Box Opening performance. Feature-map visualizations confirm that the model attends to task-relevant regions across views and perturbations, providing a direct way to inspect the visual evidence.
Figures & tables
Fig. 1: Overview of TLC-DiT. The original Multitask DiT global CLIP image feature, text feature, joint-state feature, and timestep feature are kept. The proposed path extracts frozen DINOv2 patch features, applies language-conditioned FiLM, and uses a CoordConv CNN adapter to form explicit local feature-map tokens. All conditions are concatenated for diffusion action generation. The values on the heat map range from low to high, with colors ranging from blue to red.
Component
Configuration
Transformer dimension
512
Transformer blocks
6
Attention heads
8
Dropout
0.1
Timestep embedding
256
Position encoding
RoPE, base 104
TABLE I: Diffusion Transformer configuration.
Fig. 2: Real-world tasks and evaluation milestones. Top: Teabag Putting (Lift, Insert, Close). Bottom: Match Box Opening (Push, Transfer, Pull).
Fig. 3: Local feature map comparison for TLC-DiT and five ablations. Each example shows two camera views. FiLM provides task-dependent modulation, while the CNN adapter and CoordConv produce cleaner and more spatially organized maps. The values on the heat map range from low to high, with colors ranging from blue to red.
Model
Object
Spatial
Goal
Long-10
Avg.
SmolVLA (0.45B)
89
79
88
61
79.25
Multitask DiT (0.25B)
94
86
91
75
86.50
TLC-DiT (0.29B)
98
94
95
87
93.50
(b) w/o CoordConv
95
94
93
86
92.00
(c) w/o CNN adapter
95
95
90
86
91.50
(d) w/o FiLM & CNN
97
93
95
85
92.50
TABLE II: Success Rate (%) on LIBERO
Fig. 4: Representative TLC-DiT local feature maps under LIBERO-plus perturbations. From left to right and top to bottom: light condition, background texture, object layout, and three sensor-noise examples.
Model
Camera
Robot
Language
Light
Background
Noise
Layout
Total
Multitask DiT
42.71
34.65
72.54
68.13
52.88
46.41
65.44
54.07
TLC-DiT
48.78
37.35
73.52
68.65
58.55
50.97
67.02
57.24
(b) w/o CoordConv
47.15
37.03
72.74
69.09
52.88
46.97
66.62
55.55
(c) w/o CNN adapter
45.28
38.06
71.57
67.25
54.83
49.28
66.69
55.61
(d) w/o FiLM & CNN
43.03
35.03
74.04
72.07
57.90
39.29
65.18
54.22
(e) w/o FiLM
44.84
34.65
72.54
68.21
51.86
42.54
70.95
54.53
TABLE III: Success Rate (%) under LIBERO-plus Perturbations
Fig. 5: Multi-view local feature maps during real-world execution. For each temporal phase and camera view, the columns show the RGB image, feature overlay, and local map. The values on the heat map range from low to high, with colors ranging from blue to red.
Teabag Putting
Match Box Opening
Model
Lift
Insert
Close
Push
Transfer
Pull
ACT
65
46
21
76
47
38
SmolVLA
74
73
56
12
0
0
Multitask DiT
100
50
44
80
70
52
TLC-DiT
100
92
89
76
69
55
TABLE IV: Cumulative Success Rate (%) on Real-World Tasks