datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ROMA_proactive
ROMA Proactive Streaming Dataset
Figure: Overview of ROMA's Streaming Dataset. This repository contains the Proactive subset (Green and Purple sections).
Dataset Summary
This repository contains the Proactive Interaction subset of the dataset introduced in the paper ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding.
This dataset is designed to train multimodal models for streaming video understanding, specifically focusing on tasks… See the full description on the dataset page: https://huggingface.co/datasets/EurekaTian/ROMA_proactive.ProactiveMobile
ProactiveMobile
A comprehensive, executable benchmark for proactive intelligence in mobile agents — agents that anticipate user needs and act on their own, rather than passively executing explicit commands.
📄 Paper: arXiv:2602.21858 · 🔗 Project: xiaomi-research/proactive-mobile
Overview
Each instance asks a model to infer latent user intent from four dimensions of on-device context, then produce an executable function sequence drawn from a unified function… See the full description on the dataset page: https://huggingface.co/datasets/xiaomi-research/ProactiveMobile.ProactiveMobile
ProactiveMobile
A comprehensive, executable benchmark for proactive intelligence in mobile agents — agents that anticipate user needs and act on their own, rather than passively executing explicit commands.
📄 Paper: arXiv:2602.21858 · 🔗 Project: xiaomi-research/proactive-mobile
Overview
Each instance asks a model to infer latent user intent from four dimensions of on-device context, then produce an executable function sequence drawn from a unified function… See the full description on the dataset page: https://huggingface.co/datasets/hsinv/ProactiveMobile.proactive-agent-data
Proactive Agent Dataset
Dataset for training and evaluating proactive IDE assistants.
Splits
Split
Samples
Description
train_synthetic
10489
Synthetic data from gym simulation
train_real
2476
Real IDE sessions (excluding test/val)
test
759
Test split from real sessions
val
1111
Validation split from real sessions
Fields
session_id: unique session identifier
source: "synthetic" or "real"
trigger_op_id: operation ID after which the… See the full description on the dataset page: https://huggingface.co/datasets/jiarx/proactive-agent-data.
