cs.AISep 24, 2026

Jev-Mobile: Jev as an Executor for Mobile GUI Agents

Authors: Linghua Zhang

Abstract

Vision-language models (VLMs) have become a common foundation for autonomous mobile GUI agents, but most existing systems rely on the VLM for both planning and action grounding at nearly every interaction step, leading to substantial latency and model-serving cost. We introduce Jev-Mobile, which shifts this paradigm to low-frequency VLM planning and high-frequency lightweight execution: the VLM specifies local goals, the accessibility tree defines a structured executable action space, and Jev, a fast typed decision model, repeatedly selects actions within this space. This design allows multiple GUI actions to be executed under a single VLM decision, reducing expensive VLM inference while preserving adaptive interaction. On the full AndroidWorld task suite, Jev-Mobile achieves 79% task success, compared with 78% for SeeAct-V and 84% for a Step-wise VLM baseline. Among successful trajectories, it reduces mean end-to-end execution time by 32.7% and mean model API cost by 73.4% relative to Step-wise VLM. These results show that decoupling high-level VLM reasoning from low-level action execution can substantially improve mobile GUI agent efficiency while maintaining competitive task performance.

Figures & tables

Explore similar work

CardsList
  1. MobileExplorer: Accelerating On-Device Inference for Mobile GUI Agents via Online Exploration

    May 26, 2026Runxi Huang, Liyu Zhang, Shengzhong Liu +1Graphical User Interface AgentsDevice

  2. MobileDreamer: Generative Sketch World Model for GUI Agent

    Jan 7, 2026Yilin Cao, Yufeng Zhong, Zhixiong Zeng +6Graphical User Interface AgentsDevice