Buckets:
| license: cc-by-nc-sa-4.0 | |
| language: | |
| - zh | |
| task_categories: | |
| - text-generation | |
| tags: | |
| - mobile-agent | |
| - proactive-agent | |
| - benchmark | |
| - function-calling | |
| - gui | |
| size_categories: | |
| - 1K<n<10K | |
| # ProactiveMobile | |
| A comprehensive, **executable** benchmark for **proactive intelligence in mobile agents** — agents that anticipate user needs and act on their own, rather than passively executing explicit commands. | |
| 📄 Paper: [arXiv:2602.21858](https://arxiv.org/abs/2602.21858) · 🔗 Project: [xiaomi-research/proactive-mobile](https://github.com/xiaomi-research/proactive-mobile) | |
| ## Overview | |
| Each instance asks a model to infer latent user intent from four dimensions of on-device context, then produce an **executable function sequence** drawn from a unified function pool. | |
| - **3,658** instances (Chinese, `zh`) | |
| - **Multi-answer** annotations: 1–3 target actions per instance | |
| - **61** executable APIs as the unified function pool (`function_pool.json`) | |
| - Difficulty levels: **1** (377) · **2** (1,324) · **3** (1,957) | |
| > **Note on the function pool.** For this release, semantically overlapping functions were consolidated, and the pool was pruned to the **61 APIs** used by the benchmark. | |
| ## Data Format | |
| A single JSON array (`benchmark_zh.json`). Each record: | |
| ```json | |
| { | |
| "benchmark_metadata": { | |
| "id": "deb6a1f7-7db4-4dea-b32c-bba2bae9246a", | |
| "difficulty_level": 3 | |
| }, | |
| "reference_information": { | |
| "profile": "用户是一位35岁的国际关系分析师……", | |
| "phone": "当前设备时间为晚上11点31分。手机型号为Pixel……", | |
| "world": "今天是10月10日,星期一……", | |
| "trace": [ | |
| "用户打开了 Google Scholar……", | |
| { "source": "picture", "picture": "Benchmark/aitz/android_in_the_zoo/.../xxx.png" } | |
| ] | |
| }, | |
| "recommendations": [ | |
| { | |
| "instruction": "立即打开'飞书'应用,并定位至设计团队的群聊……", | |
| "thinking": "用户在下午3点会议提醒后开启了勿扰模式……", | |
| "function": [ | |
| { | |
| "name": "view_chat_history", | |
| "parameters": { "chat_criteria": "设计团队", "app_name": "飞书" } | |
| } | |
| ] | |
| } | |
| ], | |
| "language": "zh" | |
| } | |
| ``` | |
| ### Fields | |
| | Field | Description | | |
| |---|---| | |
| | `benchmark_metadata.id` | Unique instance ID. | | |
| | `benchmark_metadata.difficulty_level` | Difficulty, 1 (easy) – 3 (hard). | | |
| | `reference_information.profile` | **User Profile** — attributes, habits, preferences. | | |
| | `reference_information.phone` | **Device Status** — time, model, battery, network, notifications. | | |
| | `reference_information.world` | **World Information** — date, weather, holidays, events. | | |
| | `reference_information.trace` | **Behavioral Trajectory** — interaction history; each step is a text description or a screenshot reference. | | |
| | `recommendations` | 1–3 target actions (the multi-answer ground truth). | | |
| | `recommendations[].instruction` | Natural-language description of the proactive action. | | |
| | `recommendations[].thinking` | Rationale linking context to the action. | | |
| | `recommendations[].function` | Executable function sequence (`name` + `parameters`) from the function pool. An empty list means "no recommendation". | | |
| > **Note on screenshots.** `trace` entries with `"source": "picture"` reference image paths (e.g. `Benchmark/aitz/...`, `Benchmark/GUI-Odyssey/...`, `Benchmark/MobileAgentBench/...`, `Benchmark/CAGUI/...`). The images are **not** included here — they come from the public **AITZ (Android in the Zoo)**, **GUI-Odyssey**, **MobileAgentBench**, and **CAGUI** datasets. Download them from their original sources and keep the relative paths to use the visual trajectories. | |
| ## Function Pool | |
| `function_pool.json` defines the **61 executable APIs**, grouped into 15 functional categories (e.g. 娱乐与媒体, 个人管理, 购物消费, 交通出行). Each entry specifies the function name, a description, and its parameter schema. The `name` values in `recommendations[].function` are drawn from this pool. | |
| ## Usage | |
| ProactiveMobile is an **evaluation benchmark** (test set). The data is a single JSON array, so loading it directly is the simplest: | |
| ```python | |
| import json | |
| data = json.load(open("benchmark_zh.json", encoding="utf-8")) | |
| print(len(data), "instances") # 3658 | |
| ``` | |
| Alternatively, with 🤗 `datasets`: | |
| ```python | |
| from datasets import load_dataset | |
| ds = load_dataset("xiaomi-research/ProactiveMobile", data_files="benchmark_zh.json", split="test") | |
| ``` | |
| ## Citation | |
| ```bibtex | |
| @article{kong2026proactivemobile, | |
| title = {ProactiveMobile: A Comprehensive Benchmark for Boosting Proactive Intelligence on Mobile Devices}, | |
| author = {Kong, Dezhi and Feng, Zhengzhao and Liang, Qiliang and others}, | |
| journal = {arXiv preprint arXiv:2602.21858}, | |
| year = {2026} | |
| } | |
| ``` | |
| ## License | |
| [CC BY-NC-SA 4.0](https://creativecommons.org/licenses/by-nc-sa/4.0/) — non-commercial use, attribution required, share-alike. | |
Xet Storage Details
- Size:
- 4.94 kB
- Xet hash:
- d283d6e48febfdbb96689a16b954b8adbde94d986a5d926237545df76f70795f
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.