cs.CLJun 22, 2026

PhoneBuddy: Training Open Models for Agentic Phone Use

Authors: Zhengyang TangXin LaiPengyuan LyuXinyuan WangTianyi BaiChenxin LiYiduo GuoHuawen Shen+18 more

Organizations: Tencent Hunyuan · Gaoling School of Artificial Intelligence, Renmin University of China · The Chinese University of Hong Kong, Shenzhen · Wuhan University

Abstract

Phones are becoming an important execution surface for general-purpose agents, but training open models for reliable phone use remains difficult because the environment that matters at deployment, real devices running real apps, is slow, stateful, side-effectful, and hard to reset or verify, while scalable mock environments only approximate real behavior. We present PhoneBuddy, a training recipe and open-model line for agentic phone use that combines a real-app environment with a mock-app environment, PhoneWorld, which reconstructs runnable mock apps from real GUI usage structure. PhoneBuddy first builds a shared supervised fine-tuning stage from trajectories collected in both environments, then compares real-app RL against mixed RL across both environments. Across a 150-task human evaluation on real phones spanning apps, mini-apps, and cross-app workflows, task success rate improves from 36.67% after supervised fine-tuning to 40.67% after real-app RL and 45.33% after mixed RL. On AndroidWorld, the same progression rises from 60.3% to 77.2% to 83.2%. These results show that mock-app training is not a replacement for real-app RL, but a complementary source of scalable, resettable, and automatically checked interaction. The gains are strongest on app and mini-app tasks, while long-horizontal cross-app workflows remain an important open challenge.

Explore similar work

CardsList
  1. PhoneWorld: Scaling Phone-Use Agent Environments

    May 28, 2026Zhengyang Tang, Yuxuan Liu, Xin Lai +21DeviceEnvironment