VLMs for Robotics

VLM: Vision-Language Model

Latest papers 392

All topics
CardsList
  1. Capability-Aware Arbitration for Semantic Intent-Based Shared Control

    Sep 21, 2026Zhaoda Du, Michael Bowman, Xiaoli ZhangUncertainty Estimation for VLA ModelsHuman-Robot Collaboration

  2. X-Planner: Event-Structured Task Planning for Embodied Intelligence

    Sep 21, 2026Howard Lu, Shalfun Li, Porter Pan +30Robot Task PlanningVLM Reasoning

  3. Think Like a World Model, Act Like a VLA: Distilling World-Model Representations into Compact Robot Policies

    Sep 21, 2026Trung Dao, Sankalp Yamsani, Jaden Park +2World ModelsRobotic Manipulation

  4. FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding

    Sep 21, 2026Hung T. Ho, Khanh D. Nguyen, Quang D. Nguyen +5VLM QuantizationEfficient VLA Model Inference

  5. StenoVLA-3D: 3D-Aware Reasoning VLA for Navigation Through Gastrointestinal Stenoses

    Sep 21, 2026Tamima Tabassum, Yiming Huang, Tianchun Wu +7EndoscopyRobotics

  6. ActiveArena: Benchmarking and Understanding Active Perception in Robotic Manipulation

    Sep 21, 2026Yibo Li, Enshen Zhou, Rui Chen +7Memory-Augmented VLMsActive Perception

  7. RoboTalk: Learning Multi-Robot Communication and Coordination from Multimodal Demonstrations

    Sep 21, 2026Dorian Benhamou Goldfajn, Mason Nakamura, Saaduddin Mahmud +3Multi-Robot SystemsLanguage-Conditioned Robot Manipulation

  8. Topology-Informed Visual Prompting For Vision Language Action Policies

    Sep 20, 2026Haoyang Wu, Abhinav Kumar, Dmitry BerensonLanguage-Conditioned Robot ManipulationVision-Language-Action Models

  9. ReVeal: A Reconstruction-Aware Real-to-Sim Framework for VLA Policy Evaluation

    Sep 20, 2026Xinyi Wang, Heng Hao, Wenjun Hu +53D ReconstructionRobotic Manipulation

  10. PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing

    Sep 20, 2026Donghao Zhou, Jia-Hui Pan, Fan Zhang +6Long-Horizon Robotic ManipulationRobotics

  11. Beyond Appearance Shifts: Task-Semantic Action Calibration for VLA Models

    Sep 20, 2026Shuaijun Liu, Feiyang You, Chengyu Wu +5Vision-Language-Action ModelsVLMs for Robotics

  12. An Empirical Study and Open Testbed for Federated Fine-Tuning of Vision-Language-Action Models

    Sep 19, 2026Zhekai Duan, Kevin Ziyang Xie, Xinyu Tan +5Vision-Language-Action ModelsVLMs for Robotics

  13. StageGuard: Learning Stage Transitions for Long-Horizon Robot Tasks via Agentic Distillation

    Sep 17, 2026Jinbang Huang, Yuanzhao Hu, Zhiyuan Li +6Long-Horizon Robotic ManipulationRobot Task Planning

  14. V2-STRep: VLM-Grounded Structured Task Representations for Reusable Robot Skills Acquired from Generated Videos

    Sep 17, 2026Yexin Hu, Dongheui LeeZero-Shot Robotic ManipulationGeneralization in Robotic Manipulation

  15. Spatial-Semantic Uncertainty in VLM-Based Target Search: Balancing Exploration and Identification

    Sep 17, 2026Alkesh K. Srivastava, Jonathan Diller, Vijay Kumar +1Belief-Space PlanningVision-Language Models

  16. Compliance for Free: Learning Identifiable Impedance via Bilateral Teleoperation

    Sep 17, 2026Harsha Guda, Adrià Colomé, Carme TorrasRobotic ControlContact-Rich Robotic Manipulation

  17. Recovering Aggressively Pruned Vision-Language-Action Models with Offline Hidden-State Distillation

    Sep 17, 2026Chiyoung Kim, Sanghyuk Roy Choi, Minhyeok LeeEfficient VLA ModelsVision-Language-Action Models

  18. VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

    Sep 17, 2026Zhongbo Zhang, Jiayi Jin, Yifan Wang +4Active PerceptionSpatial Reasoning Benchmarks

  19. FASA: Feedback-Aware Sampling Adaptation for Efficient Diffusion-Based VLA Models

    Sep 16, 2026Yuchen Han, Jianhan Wu, Xiaoyang Qu +3Diffusion SamplingEfficient VLA Model Inference

  20. In-Context Robot Learning with VLM Agents

    Sep 16, 2026Dongzhou Cheng, Taoran Yi, Ye Fang +12Robot Policy LearningICL for Robotics

  21. KINO: A Keyframe Interface for VLM Planning and Whole-Body Control in Humanoid Loco-Manipulation

    Sep 16, 2026Sitong Chen, Fatemeh Zargarbashi, Jin Cheng +2Robotic ControlRL for Robotics

  22. Calibrated Probabilistic Obstruction Reasoning with Vision-Language Models for Grasping in Clutter

    Sep 16, 2026Thanh-Tuan Tran, Ngoc-Chien Chu, Thanh Nguyen Canh +3Robotic GraspingRobotic Manipulation

  23. VLA-ULAP: Interleaving Cloud VLA Calls with Ultra-Lightweight Local Action Prediction at the Edge

    Sep 16, 2026Deyu Cao, Ryuji Oi, Kosuke Matsushima +4Efficient VLA Model InferenceVision-Language-Action Models

  24. ActiveScale: Scaling Active Perception for Robots across Model, Data, and Hardware

    Sep 16, 2026Shuai Zhou, Kaisheng Pang, Wenxuan Song +3Active Vision for Robot ManipulationRobotic Manipulation

  25. ForceDelta-VLA: Distilling Force-Conditioned ActionCorrections for Contact-Rich Manipulation

    Sep 16, 2026Ju Dong, Yu Fu, Jian Chen +9Contact-Rich Robotic ManipulationForce-Controlled Robotic Manipulation

  26. A Comprehensive Review of Generative Physical Artificial Intelligence

    Sep 16, 2026Satyam Gaba, Krutiksinh Rana, Siva Sai +2Robot Foundation ModelsDiffusion Policy

  27. Intrinsic Robot Rewarding: Reusing VLA Representations for Autonomous Evaluation and Policy Improvement

    Sep 15, 2026Tobias Schaffer, Mohab Elkhayat, Daniela Nicklas +2RL for RoboticsIntrinsic Reward Methods