Organizations: Department of Intermedia Art and Science, Waseda University, Tokyo, Japan · SB Intuitions Corp., Tokyo, Japan · Department of Electronic and Physical Systems, Waseda University, Tokyo, Japan · Department of Computer Science and Engineering, Waseda University, Tokyo, Japan · National Institute of Advanced Industrial Science and Technology (AIST), Tokyo, Japan
Language-conditioned robot policies have made clear progress in multitask manipulation, but task-relevant local visual evidence usually stays hidden inside a visual backbone or attention layers. This leaves the policy difficult to inspect and fragile under visual change, two symptoms of a missing explicit, task-aligned local visual channel. We present TLC-DiT, a plug-in extension of the Multitask Diffusion Transformer (DiT) policy that adds explicit task-guided local visual feature maps without changing the diffusion objective or the action-generation process. For each camera view, frozen DINOv2 patch features are modulated by the CLIP task embedding through FiLM and refined by a lightweight CoordConv CNN adapter into smooth spatial maps, which are concatenated with the original global image, language, joint-state, and timestep conditions. On LIBERO, TLC-DiT reaches a 93.5% average success rate, compared with 86.5% for Multitask DiT and 79.25% for SmolVLA. On LIBERO-plus, the total success rate improves from 54.07% to 57.24%, with larger gains under camera, background, and sensor-noise changes. In real-world bimanual tasks, TLC-DiT raises Teabag Putting completion from 44% to 89% while maintaining comparable Match Box Opening performance. Feature-map visualizations confirm that the model attends to task-relevant regions across views and perturbations, providing a direct way to inspect the visual evidence.
Figures & tables
Fig. 1: Overview of TLC-DiT. The original Multitask DiT global CLIP image feature, text feature, joint-state feature, and timestep feature are kept. The proposed path extracts frozen DINOv2 patch features, applies language-conditioned FiLM, and uses a CoordConv CNN adapter to form explicit local feature-map tokens. All conditions are concatenated for diffusion action generation. The values on the heat map range from low to high, with colors ranging from blue to red.
Component
Configuration
Transformer dimension
512
Transformer blocks
6
Attention heads
8
Dropout
0.1
Timestep embedding
256
Position encoding
RoPE, base 104
TABLE I: Diffusion Transformer configuration.
Fig. 2: Real-world tasks and evaluation milestones. Top: Teabag Putting (Lift, Insert, Close). Bottom: Match Box Opening (Push, Transfer, Pull).
Fig. 3: Local feature map comparison for TLC-DiT and five ablations. Each example shows two camera views. FiLM provides task-dependent modulation, while the CNN adapter and CoordConv produce cleaner and more spatially organized maps. The values on the heat map range from low to high, with colors ranging from blue to red.
Model
Object
Spatial
Goal
Long-10
Avg.
SmolVLA (0.45B)
89
79
88
61
79.25
Multitask DiT (0.25B)
94
86
91
75
86.50
TLC-DiT (0.29B)
98
94
95
87
93.50
(b) w/o CoordConv
95
94
93
86
92.00
(c) w/o CNN adapter
95
95
90
86
91.50
(d) w/o FiLM & CNN
97
93
95
85
92.50
TABLE II: Success Rate (%) on LIBERO
Fig. 4: Representative TLC-DiT local feature maps under LIBERO-plus perturbations. From left to right and top to bottom: light condition, background texture, object layout, and three sensor-noise examples.
Model
Camera
Robot
Language
Light
Background
Noise
Layout
Total
Multitask DiT
42.71
34.65
72.54
68.13
52.88
46.41
65.44
54.07
TLC-DiT
48.78
37.35
73.52
68.65
58.55
50.97
67.02
57.24
(b) w/o CoordConv
47.15
37.03
72.74
69.09
52.88
46.97
66.62
55.55
(c) w/o CNN adapter
45.28
38.06
71.57
67.25
54.83
49.28
66.69
55.61
(d) w/o FiLM & CNN
43.03
35.03
74.04
72.07
57.90
39.29
65.18
54.22
(e) w/o FiLM
44.84
34.65
72.54
68.21
51.86
42.54
70.95
54.53
TABLE III: Success Rate (%) under LIBERO-plus Perturbations
Fig. 5: Multi-view local feature maps during real-world execution. For each temporal phase and camera view, the columns show the RGB image, feature overlay, and local map. The values on the heat map range from low to high, with colors ranging from blue to red.
Teabag Putting
Match Box Opening
Model
Lift
Insert
Close
Push
Transfer
Pull
ACT
65
46
21
76
47
38
SmolVLA
74
73
56
12
0
0
Multitask DiT
100
50
44
80
70
52
TLC-DiT
100
92
89
76
69
55
TABLE IV: Cumulative Success Rate (%) on Real-World Tasks
Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the action backbone. We introduce WorldDiT, a unified diffusion transformer architecture that couples action generation with visual world modeling and achieves strong performance without a large pretrained VLM action backbone. During training, a single diffusion transformer generates continuous action chunks and predicts normalized RGB patch targets from future camera frames. Across four LIBERO simulation suites, WorldDiT lies on the reported Pareto frontier for total model parameters and mean success among methods reporting all four suites. These results provide a strong sub-billion-parameter baseline for future scaling studies.
Robot manipulation policies must generalize across visual shifts while preserving scene context relevant to action. General-purpose vision encoders are not tailored to visuomotor control, while object-centric approaches often use segmentation masks as hard filters that discard potentially useful context. We propose RoboMP-DINOv2 (Robotics Mask-Prompted DINOv2), a full-scene vision encoder that treats masks as spatial prompts rather than visibility filters. It extracts dense DINOv2 features from the full observation, injects learned region-specific embeddings at masked locations, and jointly contextualizes prompted and unprompted tokens for action prediction. We further introduce masked-region color randomization (MCR) to improve appearance robustness, yielding RoboMP-DINOv2-MCR. Across seven simulated manipulation settings, RoboMP-DINOv2 achieves 60.7% success under spatial shifts and 59.7% under scene clutter, compared with 50.7% and 41.0% for a DINOv2-based Diffusion Policy. Under unseen object colors, RoboMP-DINOv2-MCR achieves 72.5% success versus 35.1% for the strongest color-randomized baseline. Additional experiments and representation analyses show improved robustness while preserving behaviorally relevant scene information. Code is available at https://github.com/han20192019/RoboMP_DINOv2.
Han Qi, Heng Yang
Harvard School of Engineering and Applied Sciences Harvard University
Reinforcement learning (RL) for robotic manipulation often requires manually designing a dense reward function, which is difficult to tune and often fragile, or learning a reward from human demonstrations or preferences, which can be expensive. A recent line of work uses pretrained vision-language models (VLMs) as zero-shot reward models, replacing these costs with a single text prompt. However, we argue that a single global prompt is too coarse for long-horizon manipulation tasks with randomized initial conditions. The single-prompt VLM reward is near-flat for much of the trajectory, making early progress hard for the agent to detect. We propose Reinforced Micro-Task Learning (RMTL), an approach that decomposes a manipulation task into a small set of language-described micro-tasks and trains the agent to switch between them. At each step, the agent receives a multi-view VLM reward computed using the prompt of the currently active micro-task and averaged across multiple camera views to reduce the effect of view-specific occlusions. A reverse curriculum gradually exposes the agent to harder initial conditions, while a PPO worker is first trained with a fixed distance-based rule that selects the active micro-task. We then replace this rule with a learned hierarchical manager, turning rule-based phase selection into a fully learned hierarchical policy. We instantiate RMTL on the Fetch manipulation environment using three short stage-specific prompts and without additional prompt tuning. Experiments show that RMTL provides more informative reward signals than single-prompt VLM rewards, enabling faster learning. These results suggest that decomposing VLM rewards into micro-task-specific language prompts can substantially improve the scalability of language-guided reinforcement learning for robotic manipulation.