datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ProactiveBench
ProactiveBench
(ECCV 26)
Thomas De Min, Subhankar Roy, Stéphane Lathuilière, Elisa Ricci, and Massimiliano Mancini
Abstract.
Effective collaboration begins with knowing when to ask for help. For example, when trying to identify an occluded object, a human would ask someone to remove the obstruction. Can MLLMs exhibit a similar “proactive” behavior by requesting simple user interventions? To investigate this, we introduce ProactiveBench, a benchmark built from seven repurposed… See the full description on the dataset page: https://huggingface.co/datasets/tdemin16/ProactiveBench.ROMA_proactive
ROMA Proactive Streaming Dataset
Figure: Overview of ROMA's Streaming Dataset. This repository contains the Proactive subset (Green and Purple sections).
Dataset Summary
This repository contains the Proactive Interaction subset of the dataset introduced in the paper ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding.
This dataset is designed to train multimodal models for streaming video understanding, specifically focusing on tasks… See the full description on the dataset page: https://huggingface.co/datasets/EurekaTian/ROMA_proactive.trainproactive-execution-context-review-cases
Proactive Execution Context Review Cases
An original, synthetic teaching dataset for reviewing whether an AI system should prepare a next step, refresh its Context, or return a decision to a person.
What this is
Each record describes a fictional work scenario with an available Session summary, a candidate next step, and an expected review boundary. The material is intentionally small and illustrative; it is not a benchmark, model evaluation, product telemetry… See the full description on the dataset page: https://huggingface.co/datasets/ChengyiX/proactive-execution-context-review-cases.proactive-video-eval-suite
Proactive Video Evaluation Suite
This public repository contains four proactive/streaming-video evaluation collections with their original task annotations and packaged source videos. license: other is intentional: public accessibility does not replace the source-specific terms, and the annotations and upstream videos do not share one blanket license.
Contents
Directory
Rows found in the supplied source
Description
data/table2_timing_400/
400
Frozen… See the full description on the dataset page: https://huggingface.co/datasets/LCZZZZ/proactive-video-eval-suite.Qwen3.5_Proactive_SFTpersonalized-proactive-conversations
Dataset Card for "personalized-proactive-conversations"
More Information needed
proactive-ai-2000-splitproactive_image_generation
Dataset Card for "proactive_image_generation"
More Information needed
DeepSeek-R1-Distill-Data-5kproactive-ai-2000Reasoning-While-Asking-SFT-Datasetproactive-ai-dataset-2000
Dataset Card for "proactive-ai-dataset-2000"
More Information needed
ProactiveBenchThis is the benchmark associated with submission 19025 at ICLR 2026.
The benchmark is intended to be used with the proposed submission environments (see the source code).
The .jsonl files do not contain proper image paths but rather image path templates, as each .jsonl entry is a sample, and each sample corresponds to a different environment with its own images.
See the submitted code README for information about dataset downloading and preprocessing, and to re-run the evaluations.
