Training open models for agentic phone use with real-app and mock-app environments.
See the code
PhoneBuddy •
🌍 PhoneWorld •
🛠️ PhoneHarness •
🔐 PhonePrivacy •
🛡️ PhoneSafety
PhoneBuddy trains open phone-use agents that learn from both real phone execution and scalable PhoneWorld-style mock-app environments. The core result: real-app RL gives realism; mock-app RL gives resettable, verifiable interaction scale.
🧠 Open phone-use models • 📲 real-phone evaluation • 🧪 mock-app RL • ✅ verifier-backed tasks
Most mobile agents are evaluated as GUI controllers: observe a screen, tap, type, swipe, repeat. PhoneBuddy studies a training recipe for open phone-use models that can improve under real execution feedback while also benefiting from scalable mock-app supervision.
PhoneBuddy compares a shared SFT checkpoint, real-app RL, and mixed real+mock RL. The mixed recipe uses PhoneWorld-style mock apps as resettable environments with automatic verifiers, then evaluates whether this scalable signal transfers back to real-phone tasks and AndroidWorld.
PhoneWorld is the environment stack behind PhoneBuddy's mock-app training. Its public research release includes the evaluation runner, task definitions and verifiers, 120 benchmark tasks, 300 verified training tasks, and the Research Source Edition for all 34 mock Android apps.
Please follow the PhoneWorld repository's research license, APK access terms, and benchmark reporting guidance.
| Model | Status | Training Recipe | Notes |
|---|---|---|---|
| PhoneBuddy-4B | HF Model | Real+Mock RL | Main checkpoint used for the headline release. |
| PhoneBuddy-4B-RealApp | HF Model | Real-only RL | Ablation checkpoint without mock-app RL. |
| PhoneBuddy-0.8B | HF Model | Real+Mock RL | Smaller checkpoint for lightweight experiments. |
The public model release follows the Qwen-style XML tool-call format defined in the model chat_template.jinja. Dataset artifacts are not planned for public release at this stage.
| Model | Single-App | Cross-App | WeChat Mini-App | AndroidWorld | Avg. |
|---|---|---|---|---|---|
| PhoneBuddy-4B-SFT | 34.0 | 22.0 | 54.0 | 60.3 | 42.6 |
| PhoneBuddy-4B-Real | 54.0 | 20.0 | 48.0 | 77.2 | 49.8 |
| PhoneBuddy-4B-Real+Mock | 62.0 | 18.0 | 56.0 | 83.2 | 54.8 |
Takeaway. Real-app RL substantially improves over SFT. Adding mock-app RL further improves the average result, with the strongest gains on single-app tasks and AndroidWorld.
PhoneBuddy is one piece of a larger phone-agent stack: environments, training, runtime, privacy, and safety.
| Tag | Project | Links | Role |
|---|---|---|---|
| [Training] | Code · Project · Paper · 4B · 4B-RealApp · 0.8B | Trains open phone-use models with real-app RL and mock-app RL. | |
| [Environment] | 🌍 PhoneWorld | Code · APKs · Paper · 中文 Blog | Converts real GUI trajectories into scalable phone-use environments, tasks, verifiers, and rollouts. |
| [Runtime] | 🛠️ PhoneHarness | Code · Project · Dataset · Paper · 中文 Blog | Mixed-action phone-agent harness and benchmark across CLI, GUI, and MCP tools with trace-backed verification. |
| [Privacy] | 🔐 PhonePrivacy | Code · Paper · 中文 Blog | Verifiable privacy benchmark for phone-use agents. |
| [Safety] | 🛡️ PhoneSafety | Code · Paper | Safety evaluation for phone-use agents, separating safety from incapability. |
phonebuddy/
├── assets/
│ ├── figures/ # Paper and project figures
│ └── paper.pdf # Current paper snapshot
├── docs/ # Public documentation drafts
└── README.md
@misc{tang2026phonebuddytrainingopenmodels,
title={PhoneBuddy: Training Open Models for Agentic Phone Use},
author={Zhengyang Tang and Xin Lai and Pengyuan Lyu and Xinyuan Wang and Tianyi Bai and Chenxin Li and Yiduo Guo and Huawen Shen and Yuxuan Liu and Junyi Li and Zhengyao Fang and Yang Ding and Yi Zhang and Weinong Wang and Xingran Zhou and Liang Wu and Fei Tang and Sunqi Fan and Shangpin Peng and Zheng Ruan and Anran Zhang and Benyou Wang and Ji-Rong Wen and Rui Yan and Chengquan Zhang and Han Hu},
year={2026},
eprint={2606.23049},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2606.23049},
}
Made for open phone-use agents. Follow updates at phonebuddyai.github.io.
809 followers · starred Jul 2026
108 followers · starred Jun 2026
Python
40.7%
HTML
29.3%
JavaScript
25.6%
Kotlin
2.5%
CSS
1.3%
Training open models for agentic phone use with real-app and mock-app environments.
See the code
PhoneBuddy •
🌍 PhoneWorld •
🛠️ PhoneHarness •
🔐 PhonePrivacy •
🛡️ PhoneSafety
PhoneBuddy trains open phone-use agents that learn from both real phone execution and scalable PhoneWorld-style mock-app environments. The core result: real-app RL gives realism; mock-app RL gives resettable, verifiable interaction scale.
🧠 Open phone-use models • 📲 real-phone evaluation • 🧪 mock-app RL • ✅ verifier-backed tasks
Most mobile agents are evaluated as GUI controllers: observe a screen, tap, type, swipe, repeat. PhoneBuddy studies a training recipe for open phone-use models that can improve under real execution feedback while also benefiting from scalable mock-app supervision.
PhoneBuddy compares a shared SFT checkpoint, real-app RL, and mixed real+mock RL. The mixed recipe uses PhoneWorld-style mock apps as resettable environments with automatic verifiers, then evaluates whether this scalable signal transfers back to real-phone tasks and AndroidWorld.
PhoneWorld is the environment stack behind PhoneBuddy's mock-app training. Its public research release includes the evaluation runner, task definitions and verifiers, 120 benchmark tasks, 300 verified training tasks, and the Research Source Edition for all 34 mock Android apps.
Please follow the PhoneWorld repository's research license, APK access terms, and benchmark reporting guidance.
| Model | Status | Training Recipe | Notes |
|---|---|---|---|
| PhoneBuddy-4B | HF Model | Real+Mock RL | Main checkpoint used for the headline release. |
| PhoneBuddy-4B-RealApp | HF Model | Real-only RL | Ablation checkpoint without mock-app RL. |
| PhoneBuddy-0.8B | HF Model | Real+Mock RL | Smaller checkpoint for lightweight experiments. |
The public model release follows the Qwen-style XML tool-call format defined in the model chat_template.jinja. Dataset artifacts are not planned for public release at this stage.
| Model | Single-App | Cross-App | WeChat Mini-App | AndroidWorld | Avg. |
|---|---|---|---|---|---|
| PhoneBuddy-4B-SFT | 34.0 | 22.0 | 54.0 | 60.3 | 42.6 |
| PhoneBuddy-4B-Real | 54.0 | 20.0 | 48.0 | 77.2 | 49.8 |
| PhoneBuddy-4B-Real+Mock | 62.0 | 18.0 | 56.0 | 83.2 | 54.8 |
Takeaway. Real-app RL substantially improves over SFT. Adding mock-app RL further improves the average result, with the strongest gains on single-app tasks and AndroidWorld.
PhoneBuddy is one piece of a larger phone-agent stack: environments, training, runtime, privacy, and safety.
| Tag | Project | Links | Role |
|---|---|---|---|
| [Training] | Code · Project · Paper · 4B · 4B-RealApp · 0.8B | Trains open phone-use models with real-app RL and mock-app RL. | |
| [Environment] | 🌍 PhoneWorld | Code · APKs · Paper · 中文 Blog | Converts real GUI trajectories into scalable phone-use environments, tasks, verifiers, and rollouts. |
| [Runtime] | 🛠️ PhoneHarness | Code · Project · Dataset · Paper · 中文 Blog | Mixed-action phone-agent harness and benchmark across CLI, GUI, and MCP tools with trace-backed verification. |
| [Privacy] | 🔐 PhonePrivacy | Code · Paper · 中文 Blog | Verifiable privacy benchmark for phone-use agents. |
| [Safety] | 🛡️ PhoneSafety | Code · Paper | Safety evaluation for phone-use agents, separating safety from incapability. |
phonebuddy/
├── assets/
│ ├── figures/ # Paper and project figures
│ └── paper.pdf # Current paper snapshot
├── docs/ # Public documentation drafts
└── README.md
@misc{tang2026phonebuddytrainingopenmodels,
title={PhoneBuddy: Training Open Models for Agentic Phone Use},
author={Zhengyang Tang and Xin Lai and Pengyuan Lyu and Xinyuan Wang and Tianyi Bai and Chenxin Li and Yiduo Guo and Huawen Shen and Yuxuan Liu and Junyi Li and Zhengyao Fang and Yang Ding and Yi Zhang and Weinong Wang and Xingran Zhou and Liang Wu and Fei Tang and Sunqi Fan and Shangpin Peng and Zheng Ruan and Anran Zhang and Benyou Wang and Ji-Rong Wen and Rui Yan and Chengquan Zhang and Han Hu},
year={2026},
eprint={2606.23049},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2606.23049},
}
Made for open phone-use agents. Follow updates at phonebuddyai.github.io.
809 followers · starred Jul 2026
108 followers · starred Jun 2026
Python
40.7%
HTML
29.3%
JavaScript
25.6%
Kotlin
2.5%
CSS
1.3%