OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
Organizations: XPeng Motors
Abstract
Current benchmarks for graphical user interface (GUI) agents predominantly rely on static screenshots. However, real-world smartphone interaction routinely requires agents to process transient audio cues and temporal video dynamics that are tightly coupled with the moment of action. To bridge this gap, we introduce OmniGUI, the first step-level benchmark designed to evaluate GUI agents in omni-modal smartphone environments. OmniGUI provides continuous, interleaved multimodal inputs comprising static images, synchronous audio, and video clips at every action step. The dataset encompasses 709 expert-demonstrated episodes (2,579 action steps) across 29 applications, systematically annotated with objective multimodal dependency levels. Because dedicated omni-modal GUI agent frameworks are currently in their nascent stage, we select foundational omni-modal models capable of natively processing interleaved inputs to serve as agent proxies for our initial baselines. Our empirical evaluation reveals that while current models exhibit competency on visually static tasks, their action prediction performance degrades significantly in environments requiring synchronous temporal and auditory signals. Furthermore, ablation studies isolate specific operational bottlenecks, notably cross-modal interference when processing task-irrelevant environmental noise. The complete dataset, evaluation pipeline, and baseline prompts are provided in the supplementary material. Project page: https://omni-gui.github.io.
Figures & tables
| Benchmark | Platform | #Tasks | #Steps | Input Modalities | Per-Step | Action Output | Manual | ||
| Image | Video | Audio | |||||||
| Vision-Only Benchmarks | |||||||||
| AITW [ 18 ] | Android | 30,378 | 715K+ | ✓ | ✗ | ✗ | ✗ | Coordinate | ✗ |
| GUI-Odyssey [ 16 ] | Android | 7,735 | 74K+ | ✓ | ✗ | ✗ | ✗ | Coordinate | ✗ |
| AndroidWorld [ 17 ] | Android | 116 | – | ✓ | ✗ | ✗ | ✗ | Coordinate | ✓ |
| Mind2Web [ 5 ] | Web | 2,350 | 12K+ | ✓ | ✗ | ✗ | ✗ | DOM Element | ✗ |
| Action Primitive | Parameter Space | Semantics / Usage |
| Temporal / Idle | ||
| NONE (-1) | Wait or observe without taking action | |
| Positional Actions | ||
| TAP (0) | Click a UI element (button, link, icon) | |
| DOUBLE_TAP (1) | Trigger specific states (e.g., zoom, like) | |
| LONG_PRESS (2) | Open context menus or select items | |
| Model | Overall | Localization | Semantic Understand. | Cross-modal Discr. | Temporal Reason. | Instant Resp. | ||||||||||||||||||
| TM | EM | SR | GP | TM | EM | SR | GP | TM | EM | SR | GP | TM | EM | SR | GP | TM | EM | SR | GP | TM | EM | SR | GP | |
| Proprietary Models | ||||||||||||||||||||||||
| Gemini 3 Pro [ 9 ] | 80.0 | 63.6 | 33.4 | 43.6 | 86.3 | 76.2 | 55.9 | 62.6 | 77.4 | 61.1 | 31.4 | 42.0 | 76.6 | 59.1 | 30.1 | 41.3 | 78.9 | 61.0 | 22.7 | 36.9 | 81.8 | 62.7 | 27.6 | 35.6 |
| Gemini 3 Flash [ 9 ] | 78.3 | 61.3 | 30.3 | 43.5 | 85.0 | 75.6 | 53.1 | 63.1 | 75.3 | 58.5 | 25.5 | 41.1 | 72.8 | 56.0 | 23.5 | 38.7 | 80.0 | 60.3 | 25.3 | 39.4 | 79.2 | 57.9 | 22.8 | 34.2 |
| Gemini 2.5 Pro [ 4 ] | 75.7 | 44.1 | 15.5 | 26.3 | 86.1 | 58.1 | 31.7 | 41.5 | 72.8 | 37.7 | 11.7 | 22.4 | 70.6 | 40.1 | 13.2 | 25.1 | 73.8 | 44.3 | 9.7 | 22.5 | 76.6 | 42.1 | 11.0 | 19.5 |
| Gemini 2.5 Flash [ 4 ] | 69.5 | 37.8 | 12.4 | 24.5 | 75.1 | 50.9 | 29.0 | 42.6 | 70.4 | 34.3 | 8.0 | 18.2 | 64.9 | 35.7 | 11.8 | 25.3 | 67.7 | 35.1 | 9.1 | 21.8 | 71.0 | 34.5 | 3.9 | 13.7 |
| Model | Modality Input | AV-Critical (34.9%) | AV-Supportive (38.6%) | AV-Present (26.5%) | Overall (100%) | ||||||||||||
| TM | EM | SR | GP | TM | EM | SR | GP | TM | EM | SR | GP | TM | EM | SR | GP | ||
| Proprietary Models | |||||||||||||||||
| Gemini 3 Pro | Full (I+A+V) | 76.9 | 57.9 | 33.2 | 42.2 | 79.6 | 65.2 | 34.4 | 45.0 | 84.7 | 69.0 | 33.0 | 44.4 | 80.0 | 63.6 | 33.4 | 43.6 |
| No Audio (I+V) | 74.4 (-2.5) | 55.9 (-2.0) | 28.7 (-4.5) | 39.6 (-2.6) | 79.6 | 64.1 (-1.1) | 34.1 (-0.3) | 44.5 (-0.5) | 85.0 (+0.3) | 70.9 (+1.9) | 36.2 (+3.2) | 46.9 (+2.5) | 79.2 (-0.8) | 63.0 (-0.6) | 32.7 (-0.7) | 43.3 (-0.3) | |
| No Video (I+A) | 69.1 (-7.8) | 50.0 (-7.9) | 16.0 (-17.2) | 34.5 (-7.7) | 74.6 (-5.0) | 59.2 (-6.0) | 25.2 (-9.2) | 41.5 (-3.5) | 81.3 (-3.4) | 67.7 (-1.3) | 35.1 (+2.1) | 48.0 (+3.6) | 74.4 (-5.6) | 58.1 (-5.5) | 24.4 (-9.0) | 40.7 (-2.9) | |
| No AV (Img Only) | 67.3 (-9.6) | 48.9 (-9.0) | 17.2 (-16.0) | 34.5 (-7.7) | 75.3 (-4.3) | 59.0 (-6.2) | 26.3 (-8.1) | 41.0 (-4.0) | 81.8 (-2.9) | 68.9 (-0.1) | 37.8 (+4.8) | 48.8 (+4.4) | 74.2 (-5.8) | 58.0 (-5.6) | 26.1 (-7.3) | 40.8 (-2.8) | |
| Model | Instruction | AV-Critical (35.4%) | AV-Supportive (38.4%) | AV-Present (26.2%) | Overall (100%) | ||||||||||||
| TM | EM | SR | GP | TM | EM | SR | GP | TM | EM | SR | GP | TM | EM | SR | GP | ||
| Gemini 3 Pro | Text (Baseline) | 76.9 | 57.9 | 33.2 | 42.2 | 79.6 | 65.2 | 34.4 | 45.0 | 84.7 | 69.0 | 33.0 | 44.4 | 80.0 | 63.6 | 33.4 | 43.6 |
| TTS Voice | 69.0 (-7.9) | 52.1 (-5.8) | 29.1 (-4.1) | 39.4 (-2.8) | 74.5 (-5.1) | 59.9 (-5.3) | 27.7 (-6.7) | 40.3 (-4.7) | 81.6 (-3.1) | 67.3 (-1.7) | 35.8 (+2.8) | 46.4 (+2.0) | 74.3 (-5.7) | 59.1 (-4.5) | 30.3 (-3.1) | 41.6 (-2.0) | |
| Qwen3-Omni | Text (Baseline) | 58.0 | 29.4 | 7.0 | 18.5 | 65.1 | 33.5 | 4.8 | 17.4 | 66.9 | 34.4 | 3.2 | 16.0 | 63.1 | 32.3 | 5.1 | 17.4 |
| TTS Voice | 55.7 (-2.3) | 26.6 (-2.8) | 6.0 (-1.0) | 16.8 (-1.7) | 58.8 (-6.3) | 27.2 (-6.3) | 1.8 (-3.0) | 14.7 (-2.7) | 62.6 (-4.3) | 33.9 (-1.5) | 3.7 (+0.5) | 17.7 (+1.7) | 58.7 (-4.4) | 28.5 (-3.8) | 3.8 (-1.3) | 16.2 (-1.2) | |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Temperature | Max Tokens | Top_p | Top_k | Do_Sample | Seed |
| Gemini 3.0 Pro | 0.0 | 4096 | - | - | - | - |
| Gemini 3.0 Flash | 0.0 | 4096 | - | - | - | - |
| Gemini 2.5 Pro | 0.0 | 4096 | - | - | - | - |
| Gemini 2.5 Flash | 0.0 | 4096 | - | - | - | - |
| Qwen3-Omni | 0.0 | 4096 | - | - | - | - |
| MiniCPM-o 4.5 | 0.0 | 4096 | - | - | False | - |
| Chinese Applications (ZH) | English Applications (EN) | ||||||||||||||||||||
| App Name | Ep. | Stp. | Task Dimensions | Modality Dep. | App Name | Ep. | Stp. | Task Dimensions | Modality Dep. | ||||||||||||
| Loc | Sem | CrM | Tmp | Ins | Cri | Sup | Pre | Loc | Sem | CrM | Tmp | Ins | Cri | Sup | Pre | ||||||
| Bilibili | 29 | 91 | 7 | 5 | 6 | 6 | 5 | 10 | 15 | 4 | Duolingo | 35 | 133 | 5 | 5 | 5 | 15 | 5 | 16 | 11 | 8 |
| Douyin | 29 | 97 | 7 | 5 | 6 | 6 | 5 | 12 | 8 | 9 | Vimeo | 25 | 86 | 5 | 5 | 5 | 5 | 5 | 12 | 6 | 7 |
| Meituan | 26 | 103 | 6 | 5 | 6 | 5 | 4 | 4 | 5 | 17 | TED | 25 | 115 | 5 | 5 | 5 | 5 | 5 | 16 | 6 | 3 |
| QQ Music | 25 | 95 | 5 | 5 | 5 | 5 | 5 | 1 | 11 | 13 | Snapchat | 25 | 104 | 5 | 5 | 5 | 5 | 5 | 8 | 11 | 6 |