VLMs for Robotics

VLM: Vision-Language Model

Latest papers 392

All topics
CardsList
  1. dVLA-RL: Reinforcement Learning over Denoising Trajectories for Discrete Diffusion Vision-Language-Action Models

    Jun 22, 2026Yuhao Wu, Yitian Liu, Weijie Shen +13Vision-Language-Action ModelsVLMs for Robotics

  2. KEMO: Event-Driven Keyframe Memory for Long-Horizon Robot Manipulation with VLA Policies

    Jun 22, 2026Yihan Zeng, Minghao Ye, Yiyuan Chen +4Long-Horizon Robotic ManipulationBimanual Robotic Manipulation

  3. Bridging Semantics and Kinematics: A Modular Framework for Zero-Shot Robotic Manipulation

    Jun 22, 2026Ali Alabbas, Dipshikha Das, Camillo Murgia +3Zero-Shot Robotic ManipulationLanguage-Conditioned Robot Manipulation

  4. UniFS: Unified Fast-to-Slow Hierarchical Architecture for Vision-Language-Action Models

    Jun 22, 2026Lin Sun, Zhiwei Guan, Conglin Wang +7Efficient VLM InferenceEfficient VLA Model Inference

  5. EgoSteer: An Open-Source Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos

    Jun 21, 2026Yifan Zhong, Zhang Chen, Tianrui Guan +15Long-Horizon Robotic ManipulationRobot Skill Learning

  6. PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models

    Jun 21, 2026Xianghui Wang, Feng Chen, Wenbo Zhang +4Efficient VLA ModelsVision-Language-Action Models

  7. Robot Critics that Sweat the Small Stuff

    Jun 19, 2026Sruthi Sudhakar, Junbang Liang, Sreehari Rammohan +3Robot Failure DetectionVLM Adaptation

  8. Slow Brain, Fast Planner: Latency-Resilient VLM-Augmented Urban Navigation

    Jun 18, 2026Zhenghao "Mark'' Peng, Honglin He, Quanyi Li +2Trajectory GenerationRobot Navigation

  9. ReSiReg: Towards Spatially Consistent Semantics in Language-Conditioned Robotic Tasks

    Jun 17, 2026Simon Schwaiger, David Seyser, Alessandro Scherl +2Vision-Language ModelsVision-Language Grounding

  10. EffiNav: Fusing Depth and Vision-Language for Efficient Object Goal Navigation

    Jun 17, 2026Zecheng Yin, Benedict Jun MaVLMs for RoboticsObject-Goal Navigation

  11. Guava: Distilling Frontier VLMs into a Compact Agent through a Robotic Manipulation Harness

    Jun 16, 2026Haowen Liu, Xirui Li, Shaoxiong Yao +5Long-Horizon Robotic ManipulationVLM Distillation

  12. PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space

    Jun 16, 2026Bochen Yang, Lianlei ShanEfficient VLA ModelsLatent Visual Reasoning

  13. GeneralVLA-2: Geometry-Aware Reconstruction and Governed Memory for Robot Planning

    Jun 16, 2026Haoyu Wang, Guoqing Ma, Zeyu Zhang +3LLM Agent MemoryMulti-View 3D Reconstruction

  14. ROVE: Unlocking Human Interventions for Humanoid Manipulation via Reinforcement Learning

    Jun 15, 2026Wei Xiao, Weiliang Tang, Yuying Ge +4RL for RoboticsRobotic Manipulation

  15. CrossMaps: Confidence-Aware Open-Vocabulary Semantic Mapping for Rover Navigation

    Jun 15, 2026Jan-Niklas Klein, Sona Ghahremani, Christian Medeiros Adriano +1Robot NavigationSemantic Mapping

  16. Decoupled Object-Centric Video Understanding for Generating Robotic Manipulation Commands

    Jun 15, 2026Thanh Nguyen Canh, Thanh-Tuan Tran, Haolan Zhang +3Video Action RecognitionRobotic Manipulation

  17. SimWeaver: Zero-Shot RGB Sim-to-Real for Deformable Manipulation

    Jun 13, 2026Wenkang Hu, Haoran Wang, Yitong Li +10Zero-Shot Robotic ManipulationRobot Skill Learning

  18. RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation

    Jun 11, 2026Dayu Xia, Yue Shi, Yao Mu +7VLM EvaluationRobotic Manipulation

  19. Bounding Boxes as Goals: Language-Conditioned Grasping via Neuro-Symbolic Planning

    Jun 11, 2026Allison Andreyev, Landon Eum, Nestor Tiglao +1Robotic GraspingZero-Shot Robotic Manipulation

  20. AIR-VLA+: Decoupling Movement and Manipulation via Cascaded Dual-Action Decoders with Asymmetric MoE for Aerial Robots

    Jun 11, 2026Jianli Sun, Bin Tian, Qiyao Zhang +6VLMs for RoboticsAerial Manipulation

  21. DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?

    Jun 10, 2026Jadelynn Dao, Milan Ganai, Yasmina Abukhadra +7Test-Time Scaling for VLMsInference-Time Scaling

  22. Learning What to Say to Your VLA: Mostly Harmless Vision Language Action Model Steering

    Jun 10, 2026Hyun Joe Jeong, Gokul Swamy, Andrea BajcsyLanguage Model-Based ControlVision-Language-Action Models