datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ProactiveVideoQA
ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
📄 arXiv Paper |
🖥️ Github Code |
📦 Data
Introduction
ProactiveVideoQA is the first comprehensive benchmark designed to evaluate a system's ability to engage in proactive interaction in multimodal dialogue settings.
Unlike traditional turn-by-turn dialogue systems, in proactive intraction model need to determine when to repsond during… See the full description on the dataset page: https://huggingface.co/datasets/wangyueqian/ProactiveVideoQA.ProactiveBench
ProactiveBench
(ECCV 26)
Thomas De Min, Subhankar Roy, Stéphane Lathuilière, Elisa Ricci, and Massimiliano Mancini
Abstract.
Effective collaboration begins with knowing when to ask for help. For example, when trying to identify an occluded object, a human would ask someone to remove the obstruction. Can MLLMs exhibit a similar “proactive” behavior by requesting simple user interventions? To investigate this, we introduce ProactiveBench, a benchmark built from seven repurposed… See the full description on the dataset page: https://huggingface.co/datasets/tdemin16/ProactiveBench.ROMA_proactive
ROMA Proactive Streaming Dataset
Figure: Overview of ROMA's Streaming Dataset. This repository contains the Proactive subset (Green and Purple sections).
Dataset Summary
This repository contains the Proactive Interaction subset of the dataset introduced in the paper ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding.
This dataset is designed to train multimodal models for streaming video understanding, specifically focusing on tasks… See the full description on the dataset page: https://huggingface.co/datasets/EurekaTian/ROMA_proactive.Proactive-CSI-ProcessedtrainProactiveMobile
ProactiveMobile
A comprehensive, executable benchmark for proactive intelligence in mobile agents — agents that anticipate user needs and act on their own, rather than passively executing explicit commands.
📄 Paper: arXiv:2602.21858 · 🔗 Project: xiaomi-research/proactive-mobile
Overview
Each instance asks a model to infer latent user intent from four dimensions of on-device context, then produce an executable function sequence drawn from a unified function… See the full description on the dataset page: https://huggingface.co/datasets/xiaomi-research/ProactiveMobile.proactive-execution-context-review-cases
Proactive Execution Context Review Cases
An original, synthetic teaching dataset for reviewing whether an AI system should prepare a next step, refresh its Context, or return a decision to a person.
What this is
Each record describes a fictional work scenario with an available Session summary, a candidate next step, and an expected review boundary. The material is intentionally small and illustrative; it is not a benchmark, model evaluation, product telemetry… See the full description on the dataset page: https://huggingface.co/datasets/ChengyiX/proactive-execution-context-review-cases.VisualActBenchproactive-video-eval-suite
Proactive Video Evaluation Suite
This public repository contains four proactive/streaming-video evaluation collections with their original task annotations and packaged source videos. license: other is intentional: public accessibility does not replace the source-specific terms, and the annotations and upstream videos do not share one blanket license.
Contents
Directory
Rows found in the supplied source
Description
data/table2_timing_400/
400
Frozen… See the full description on the dataset page: https://huggingface.co/datasets/LCZZZZ/proactive-video-eval-suite.Proactive-Sound-Effect-BenchmarkQwen3.5_Proactive_SFTProactiveMobile
ProactiveMobile
A comprehensive, executable benchmark for proactive intelligence in mobile agents — agents that anticipate user needs and act on their own, rather than passively executing explicit commands.
📄 Paper: arXiv:2602.21858 · 🔗 Project: xiaomi-research/proactive-mobile
Overview
Each instance asks a model to infer latent user intent from four dimensions of on-device context, then produce an executable function sequence drawn from a unified function… See the full description on the dataset page: https://huggingface.co/datasets/hsinv/ProactiveMobile.personalized-proactive-conversations
Dataset Card for "personalized-proactive-conversations"
More Information needed
proactive-ai-2000-splitproactive_image_generation
Dataset Card for "proactive_image_generation"
More Information needed
DeepSeek-R1-Distill-Data-5kvideo_2proactive-ai-2000Reasoning-While-Asking-SFT-Datasetproactive-ai-dataset-2000
Dataset Card for "proactive-ai-dataset-2000"
More Information needed
proactive-hoi-san-mediaProactiveBenchThis is the benchmark associated with submission 19025 at ICLR 2026.
The benchmark is intended to be used with the proposed submission environments (see the source code).
The .jsonl files do not contain proper image paths but rather image path templates, as each .jsonl entry is a sample, and each sample corresponds to a different environment with its own images.
See the submitted code README for information about dataset downloading and preprocessing, and to re-run the evaluations.
proactive-agent-data
Proactive Agent Dataset
Dataset for training and evaluating proactive IDE assistants.
Splits
Split
Samples
Description
train_synthetic
10489
Synthetic data from gym simulation
train_real
2476
Real IDE sessions (excluding test/val)
test
759
Test split from real sessions
val
1111
Validation split from real sessions
Fields
session_id: unique session identifier
source: "synthetic" or "real"
trigger_op_id: operation ID after which the… See the full description on the dataset page: https://huggingface.co/datasets/jiarx/proactive-agent-data.proactive-sample
