From Language to Motion: Task-Conditioned Focal-Stack Trajectory Integration for Microscopic Robots
Authors: Junjie Xie, Chuxuan He, Junkai Huang, Heng Zhang, Angen Ye, Yujia Song, Yuqing Li, Pengsong Zhang, +1 more
Organizations: Institute of Automation, Chinese Academy of Sciences · Zhejiang Gongshang University · Safe AI Lab, Carnegie Mellon University · Department of Mechanical and Industrial Engineering, University of Toronto, Toronto, Ontario, M5S 3G8, Canada
Microscopic robots require accurate task geometry despite changes in language, parts, and focus. We present a semantic-to-physical framework that maps instructions to constrained geometric operators, reuses frozen open-vocabulary perception, and integrates locally reliable focal-plane trajectories by confidence weighting and dynamic programming. Calibrated multi-view geometry connects 2-D paths to physical execution. Prompt, unseen-part, and geometry reconfiguration tests yield 6.30-6.59-pixel RMSE. Relative to part-specific U-Net training with 20-100 labels, the proposed zero-new-label configuration takes 15 rather than 72-165 min. Across nine part-illumination conditions, trajectory-space integration reduces RMSE from 14.41 to 6.28 pixels (56.4%) and P95 error from 20.07 to 8.13 pixels (59.5%) compared with image-first multi-focus fusion. An ablation isolates the roles of confidence and path-wise selection. In representative robot experiments, target-region coverage improves from 83.5% to 92.9%. Dispensing provides a measurable physical trace, not a task-specific limitation of the method.
Figures & tables
Fig. 1: System overview. (A) CAD-rendered examples illustrate task-level geometric reconfiguration through language prompts. (B) Trajectory candidates and local evidence are extracted at individual focal planes before integration in path space. (C) Calibrated multi-view geometry bridges the recovered 2-D paths to physical execution on the microscopic robot.
Fig. 2: Semantic-to-geometry reconfiguration. (a) Traditional part-specific retraining versus a prompt-driven workflow. (b) Natural-language instructions are converted into constrained task specifications. (c) Reusable frozen detection/segmentation models and lightweight geometry operators recover 2-D task geometry. (d) The same visual backbone serves multiple geometry requests without task-specific detector retraining.
Fig. 3: Confidence-weighted trajectory-level multi-focal integration. (a) Local trajectories and reliability are estimated separately for each focal plane. (b) Image-first and trajectory-first processing paths. (c) Confidence-weighted hypotheses and dynamic-programming selection in ordered path space. (d) The final 2-D trajectory preserves focal-source information.
Fig. 4: Real-system physical validation. (a) Microscopic manipulation platform and auxiliary setup. (b) Complementary microscope views. (c) Robot–part interaction during execution. (d) Deposited physical traces with local image-software dimensional annotations. Dispensing is used as a trace-preserving measurement task; target-relative performance is summarized separately in Table V .
Condition
RMSE (px) ↓
P95 (px) ↓
Canonical wording
6.30±0.03
8.08±0.03
Formal paraphrase
6.34±0.03
8.12±0.04
Semantic paraphrase
6.43±0.03
8.25±0.03
Colloquial wording
6.57±0.03
8.37±0.06
Unseen part D
6.39±0.03
8.19±0.03
Unseen part E
6.50±0.03
8.26±0.04
TABLE I: Language and task reconfiguration.
Method
Labels
RMSE
P95
Adapt. time
U-Net (100)
100
5.47
7.15
165 min
U-Net (50)
50
5.78
7.52
119 min
U-Net (20)
20
6.46
8.26
72 min
Ours
0
6.28
8.13
15 min
TABLE II: Task-specific supervision versus reconfiguration cost.
Method
RMSE
P95
Mean
Max
Image-first MFIF
14.41±0.04
20.07±0.09
13.50±0.07
20.09±0.06
Ours
6.28±0.04
8.13±0.05
5.93±0.05
9.25±0.05
Reduction
56.4%
59.5%
56.1%
54.0%
TABLE III: Measured trajectory error across nine part–illumination conditions.
School of Computing and Smart Systems Institute, National University of Singapore, Singapore · Department of Electrical and Computer Engineering, University of California, Los Angeles, Los Angeles, CA, USA