cs.ROAug 10, 2026

Mamba-based Selective State Space Modeling Improves the Accuracy-Complexity Tradeoff of SmolVLA Vision-Language-Action Experts

Authors: Farida Mohsen, Thowayba Elkaffash, Mohammad Reza Chalak Qazani, Mohamed Mabrok, Nader Meskin, Ali Safa

Organizations: College of Science and Engineering, Hamad Bin Khalifa University, Doha, Qatar · College of Engineering, Qatar University, Doha, Qatar · College of Science and Engineering, James Cook University, Townsville, QLD, 4814, Australia

Abstract

Vision-language-action (VLA) models face a crucial tradeoff between their task success rate and the policy-call frequency. Executing a single action per inference (N=1N=1) enables accurate robot control but comes at the cost of huge compute time overheads, making real-time implementation infeasible. On the other hand, executing longer action horizons before replanning (N≫1N\gg1) reduces compute complexity, but inevitably degrades the system's success rate. In order to improve the VLA accuracy-complexity tradeoff, this paper investigates Mamba's selective state-space modeling as an alternative to causal self-attention within the action expert of the popular SmolVLA model, widely used as a reference model for its highly accurate yet low complexity nature. We evaluate both the Mamba- and Transformer-based experts on the widely-adopted LIBERO benchmark suites across three execution horizons N ⁣∈ ⁣{1,25,50}N\!\in\!\{1,25,50\}, respectively corresponding to high, moderate and low compute complexities. Our results remarkably show that the advantage of the Mamba expert increases with the execution horizon, indicating significant success retention under long execution horizons N=50N = 50 and N=25N = 25. When N=50N = 50 actions are executed before replanning (i.e., corresponding to feasible real-time deployment), the Mamba expert outperforms the Transformer baseline by 7.8%7.8\%. In addition, when N=25N = 25 actions are executed before replanning, our Mamba expert outperforms the Transformer baseline by 3.7%3.7\%. Finally, under per-action replanning (N=1N=1), our Mamba variant matches the Transformer-based mean success rate while significantly reducing the overall model parameter complexity by 24%24\% thanks to Mamba's compute-efficient nature.

Explore similar work

CardsList
  1. SMILE: Smooth Motion for Improved Long-Horizon VLA Execution

    Aug 29, 2026Jongwoo Park, E-Ro Nguyen, Kanchana Ranasinghe +3Vision-Language-Action ModelsRobot Motion Generation

  2. Per-Group Error, Not Total MSE: Fine-Tuning Vision-Language-Action Models for 11-DoF Mobile Manipulation

    May 29, 2026Pau Montagut Bofi, Mario García Blasco, Tessa Pulli +1Mobile ManipulationVLM Adaptation

  3. Remember What You Did: Action-History Memory with Dual-Expert Denoising for Long-Horizon Vision-Language-Action Policies

    Sep 29, 2026Yaxin Zhao, Dianye Huang, Chenwei Wang +2Memory-Augmented VLMsLong-Horizon Robotic Manipulation