CoolFace
Datasetpublic

hsinv/ProactiveMobile

ProactiveMobile A comprehensive, executable benchmark for proactive intelligence in mobile agents — agents that anticipate user needs and act on their own, rather than passively executing explicit commands. 📄 Paper: arXiv:2602.21858 · 🔗 Project: xiaomi-research/proactive-mobile Overview Each instance asks a model to infer latent user intent from four dimensions of on-device context, then produce an executable function sequence drawn from a unified function… See the full description on the dataset page: https://huggingface.co/datasets/hsinv/ProactiveMobile.

sourceHugging Facecc-by-nc-sa-4.0updated 2mo agoView on Hugging Face
0likes26downloads
Dataset Card

ProactiveMobile

A comprehensive, executable benchmark for proactive intelligence in mobile agents — agents that anticipate user needs and act on their own, rather than passively executing explicit commands.

📄 Paper: arXiv:2602.21858 · 🔗 Project: xiaomi-research/proactive-mobile

Overview

Each instance asks a model to infer latent user intent from four dimensions of on-device context, then produce an executable function sequence drawn from a unified function pool.

  • —3,658 instances (Chinese, zh)
  • —Multi-answer annotations: 1–3 target actions per instance
  • —61 executable APIs as the unified function pool (function_pool.json)
  • —Difficulty levels: 1 (377) · 2 (1,324) · 3 (1,957)
Note on the function pool. For this release, semantically overlapping functions were consolidated, and the pool was pruned to the 61 APIs used by the benchmark.

Data Format

A single JSON array (benchmark_zh.json). Each record:

json
{
  "benchmark_metadata": {
    "id": "deb6a1f7-7db4-4dea-b32c-bba2bae9246a",
    "difficulty_level": 3
  },
  "reference_information": {
    "profile": "用户是一位35岁的国际关系分析师……",
    "phone":   "当前设备时间为晚上11点31分。手机型号为Pixel……",
    "world":   "今天是10月10日,星期一……",
    "trace": [
      "用户打开了 Google Scholar……",
      { "source": "picture", "picture": "Benchmark/aitz/android_in_the_zoo/.../xxx.png" }
    ]
  },
  "recommendations": [
    {
      "instruction": "立即打开'飞书'应用,并定位至设计团队的群聊……",
      "thinking": "用户在下午3点会议提醒后开启了勿扰模式……",
      "function": [
        {
          "name": "view_chat_history",
          "parameters": { "chat_criteria": "设计团队", "app_name": "飞书" }
        }
      ]
    }
  ],
  "language": "zh"
}

Fields

FieldDescription
benchmark_metadata.idUnique instance ID.
benchmark_metadata.difficulty_levelDifficulty, 1 (easy) – 3 (hard).
reference_information.profileUser Profile — attributes, habits, preferences.
reference_information.phoneDevice Status — time, model, battery, network, notifications.
reference_information.worldWorld Information — date, weather, holidays, events.
reference_information.traceBehavioral Trajectory — interaction history; each step is a text description or a screenshot reference.
recommendations1–3 target actions (the multi-answer ground truth).
recommendations[].instructionNatural-language description of the proactive action.
recommendations[].thinkingRationale linking context to the action.
recommendations[].functionExecutable function sequence (name + parameters) from the function pool. An empty list means "no recommendation".
Note on screenshots. trace entries with "source": "picture" reference image paths (e.g. Benchmark/aitz/..., Benchmark/GUI-Odyssey/..., Benchmark/MobileAgentBench/..., Benchmark/CAGUI/...). The images are not included here — they come from the public AITZ (Android in the Zoo), GUI-Odyssey, MobileAgentBench, and CAGUI datasets. Download them from their original sources and keep the relative paths to use the visual trajectories.

Function Pool

function_pool.json defines the 61 executable APIs, grouped into 15 functional categories (e.g. 娱乐与媒体, 个人管理, 购物消费, 交通出行). Each entry specifies the function name, a description, and its parameter schema. The name values in recommendations[].function are drawn from this pool.

Usage

ProactiveMobile is an evaluation benchmark (test set). The data is a single JSON array, so loading it directly is the simplest:

python
import json

data = json.load(open("benchmark_zh.json", encoding="utf-8"))
print(len(data), "instances")  # 3658

Alternatively, with 🤗 datasets:

python
from datasets import load_dataset

ds = load_dataset("xiaomi-research/ProactiveMobile", data_files="benchmark_zh.json", split="test")

Citation

bibtex
@article{kong2026proactivemobile,
  title   = {ProactiveMobile: A Comprehensive Benchmark for Boosting Proactive Intelligence on Mobile Devices},
  author  = {Kong, Dezhi and Feng, Zhengzhao and Liang, Qiliang and others},
  journal = {arXiv preprint arXiv:2602.21858},
  year    = {2026}
}

License

CC BY-NC-SA 4.0 — non-commercial use, attribution required, share-alike.