cs.ROSep 29, 2026

Video2STL: Grounding VLM-Generated Temporal Specifications for Robot Learning

Authors: Merve Atasever, Keyan Azbijari, Cagan Bakirci, Bo-Ruei Huang, Tolga Izdas, Zahra Shahrooei, Richard Yang, Erdem Biyik, +1 more

Organizations: Department of Computer Science, University of Southern California Los Angeles, CA, USA · University of Florida Gainesville, FL, USA

Abstract

Video-based policy learning is particularly promising, as it illustrates target behaviors without requiring action annotations or embodiment-matched demonstrations. A central challenge is deciding what information should be transferred from the video to the robot. Existing approaches commonly convert visual observations into scalar similarity or value signals, or ask foundation models to directly generate reward code. These approaches can make the temporal structure of a task difficult to inspect, ground, and reuse. We present Video2STL, a framework that converts observation-only videos into parametric Signal Temporal Logic (STL) specifications and uses the resulting formal representation for robot learning. A vision-language model extracts an embodiment-independent semantic event trace and constructs a bank of symbolic temporal specifications. The model determines the task structure, while numerical predicate thresholds and temporal bounds are grounded from successful robot trajectories. For policy learning, we separate short- and long-timescale temporal information: short-horizon specifications provide dense rewards through rolling-window quantitative robustness, while a causal monitor over a retained long-horizon specification provides one-time progress rewards for valid temporal prefixes. The same representation supports cross-embodiment transfer from human or animal videos to robot control. Across four manipulation tasks, Video2STL achieves 85.8%85.8\% average success-once and 67.0%67.0\% success-at-end, compared with 81.5%/59.5%81.5\%/59.5\% for native dense PPO and 65.0%/42.3%65.0\%/42.3\% for Text2Reward; in quadruped locomotion, Qwen-3.8 and GPT-5.6-based Video2STL policies achieve 100%100\% success across velocities from 0.30.3 to 2.1 m/s2.1\,\mathrm{m/s} while remaining competitive in high-speed energy efficiency. Project webpage: video2stl.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. STeP: Signal Temporal Logic for Precise Specifications for Action Generation with Vision Language Models

    Jul 20, 2026Kasra Torshizi, Anukriti Singh, Sidharth Mathur +3Signal Temporal LogicAction Generation

  2. Learning Gait-Aware Quadruped Locomotion with Temporal Logic Specifications

    Jul 1, 2026Merve Atasever, Cagan Bakirci, Alfredo Reina Corona +2LocomotionGait Dynamics

  3. WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos

    Jul 13, 2026Jiahao Liu, Yupeng Zheng, Zhongpu Xia +10Latent Actions