From Language to Motion: Task-Conditioned Focal-Stack Trajectory Integration for Microscopic Robots
Authors: Junjie Xie, Chuxuan He, Junkai Huang, Heng Zhang, Angen Ye, Yujia Song, Yuqing Li, Pengsong Zhang, +1 more
Organizations: Institute of Automation, Chinese Academy of Sciences · Zhejiang Gongshang University · Safe AI Lab, Carnegie Mellon University · Department of Mechanical and Industrial Engineering, University of Toronto, Toronto, Ontario, M5S 3G8, Canada
Microscopic robots require accurate task geometry despite changes in language, parts, and focus. We present a semantic-to-physical framework that maps instructions to constrained geometric operators, reuses frozen open-vocabulary perception, and integrates locally reliable focal-plane trajectories by confidence weighting and dynamic programming. Calibrated multi-view geometry connects 2-D paths to physical execution. Prompt, unseen-part, and geometry reconfiguration tests yield 6.30-6.59-pixel RMSE. Relative to part-specific U-Net training with 20-100 labels, the proposed zero-new-label configuration takes 15 rather than 72-165 min. Across nine part-illumination conditions, trajectory-space integration reduces RMSE from 14.41 to 6.28 pixels (56.4%) and P95 error from 20.07 to 8.13 pixels (59.5%) compared with image-first multi-focus fusion. An ablation isolates the roles of confidence and path-wise selection. In representative robot experiments, target-region coverage improves from 83.5% to 92.9%. Dispensing provides a measurable physical trace, not a task-specific limitation of the method.
Figures & tables
Fig. 1: System overview. (A) CAD-rendered examples illustrate task-level geometric reconfiguration through language prompts. (B) Trajectory candidates and local evidence are extracted at individual focal planes before integration in path space. (C) Calibrated multi-view geometry bridges the recovered 2-D paths to physical execution on the microscopic robot.
Fig. 2: Semantic-to-geometry reconfiguration. (a) Traditional part-specific retraining versus a prompt-driven workflow. (b) Natural-language instructions are converted into constrained task specifications. (c) Reusable frozen detection/segmentation models and lightweight geometry operators recover 2-D task geometry. (d) The same visual backbone serves multiple geometry requests without task-specific detector retraining.
Fig. 3: Confidence-weighted trajectory-level multi-focal integration. (a) Local trajectories and reliability are estimated separately for each focal plane. (b) Image-first and trajectory-first processing paths. (c) Confidence-weighted hypotheses and dynamic-programming selection in ordered path space. (d) The final 2-D trajectory preserves focal-source information.
Fig. 4: Real-system physical validation. (a) Microscopic manipulation platform and auxiliary setup. (b) Complementary microscope views. (c) Robot–part interaction during execution. (d) Deposited physical traces with local image-software dimensional annotations. Dispensing is used as a trace-preserving measurement task; target-relative performance is summarized separately in Table V .
Condition
RMSE (px) ↓
P95 (px) ↓
Canonical wording
6.30±0.03
8.08±0.03
Formal paraphrase
6.34±0.03
8.12±0.04
Semantic paraphrase
6.43±0.03
8.25±0.03
Colloquial wording
6.57±0.03
8.37±0.06
Unseen part D
6.39±0.03
8.19±0.03
Unseen part E
6.50±0.03
8.26±0.04
TABLE I: Language and task reconfiguration.
Method
Labels
RMSE
P95
Adapt. time
U-Net (100)
100
5.47
7.15
165 min
U-Net (50)
50
5.78
7.52
119 min
U-Net (20)
20
6.46
8.26
72 min
Ours
0
6.28
8.13
15 min
TABLE II: Task-specific supervision versus reconfiguration cost.
Method
RMSE
P95
Mean
Max
Image-first MFIF
14.41±0.04
20.07±0.09
13.50±0.07
20.09±0.06
Ours
6.28±0.04
8.13±0.05
5.93±0.05
9.25±0.05
Reduction
56.4%
59.5%
56.1%
54.0%
TABLE III: Measured trajectory error across nine part–illumination conditions.
This paper presents a modular training-free framework for zero-shot, language-guided robotic manipulation in semi-structured environments. The architecture bridges the gap between high-level reasoning and low-level kinematics by decomposing the vision-action pipeline into three stages: visual perception, semantic interpretation, and task execution. To overcome the spatial ambiguity and semantic hallucinations inherent in standard Vision-Language Models (VLMs), the perception module employs FastSAM and Set-of-Mark (SoM) prompting to dynamically generate grounded, alphanumeric visual anchors. The same foundation model then operates purely as a Large Language Model (LLM) to act as a semantic router, translating unconstrained human directives into verifiable, reconfigurable configurations. Finally, these configurations are dynamically parsed by a Task Orchestrator into MoveIt Task Constructor (MTC) to generate collision-free trajectories. The framework is evaluated across two zero-shot experimental setups: unconstrained open-world sequential manipulation and dense relational spatial reasoning, achieving a 62% end-to-end task success rate across both scenarios, demonstrating its capacity to reliably execute complex physical actions without domain-specific training or manual coordinate programming.
Ali Alabbas, Dipshikha Das, Camillo Murgia +3
Department of Electronic, Software and Advanced Manufacturing Engineering, Atlantic Technological University, Galway, Ireland · Atlantic Technological University, Galway, Ireland
Understanding natural human instructions is crucial for deploying robots in human-centric environments. We study multimodal instruction grounding, where language and gesture provide complementary but uncertain cues. We present MIGU, a modular framework that combines semantic and geometric evidence into a unified grounding belief and connects it to manipulation planning. MIGU constructs a 3D geometric likelihood by propagating viewing-direction and depth uncertainty through eye-finger geometry while accounting for hand-direction estimation error. A vision-language model (VLM) provides semantic priors over candidate objects and regions, which are combined with the geometric likelihood through Bayes-inspired fusion. The resulting belief supports behavior planning to either proceed directly to downstream planning or request clarification. Grounded targets then define goals for mobile manipulation and tabletop task-and-motion planning. On a real-world benchmark, MIGU outperforms all evaluated baselines, while ablations support the benefit of explicit multimodal uncertainty modeling. Project website: multimodal-instruction.github.io
Mingke Lu, Anxing Xiao, David Hsu
School of Computing and Smart Systems Institute, National University of Singapore, Singapore · Department of Electrical and Computer Engineering, University of California, Los Angeles, Los Angeles, CA, USA
Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently. Yet, today's best systems depend on depth sensors, multi-camera rigs, or pre-built maps, limiting the hardware they support and increasing deployment cost. We introduce Robostral Navigate, an 8B vision-language model built around this scalability objective. The model consumes only a stream of monocular RGB images - the most ubiquitous sensor across robotic platforms and predicts waypoints by pointing to the next target location in the current camera view. Operating purely in image space, rather than robot-specific coordinates, makes the policy naturally robust to changes in camera intrinsics and scene scale, enabling deployment across wheeled, legged, and aerial robots without recalibration. We generate 2.4 million trajectories across 350k simulated scenes to reduce the reliance on real-world data collection and scale easily. We further introduce a prefix-caching training recipe that packs entire episodes into single training sequences, reducing training tokens by 22x and cutting training time from months to days. A tree-based attention mask prevents conditioning on previous ground-truth actions, encouraging visually grounded action prediction, and reinforcement learning is used to further improve exploration and recovery capabilities. On the Room-to-Room and Room-Across-Room in Continuous Environments (R2R-CE and RxR-CE) benchmarks, Robostral Navigate sets a new state of the art. On R2R-CE, it achieves a 77.4% success rate, surpassing the best monocular method by 10.5 points and the strongest depth- or multi-camera system by 5.3 points despite using only a single RGB camera. On RxR-CE, it reaches 75.1% success rate, outperforming all monocular baselines.