proactive
ProactiveVideoQA
ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
📄 arXiv Paper |
🖥️ Github Code |
📦 Data
Introduction
ProactiveVideoQA is the first comprehensive benchmark designed to evaluate a system's ability to engage in proactive interaction in multimodal dialogue settings.
Unlike traditional turn-by-turn dialogue systems, in proactive intraction model need to determine when to repsond during… See the full description on the dataset page: https://huggingface.co/datasets/wangyueqian/ProactiveVideoQA.ProactiveBench
ProactiveBench
(ECCV 26)
Thomas De Min, Subhankar Roy, Stéphane Lathuilière, Elisa Ricci, and Massimiliano Mancini
Abstract.
Effective collaboration begins with knowing when to ask for help. For example, when trying to identify an occluded object, a human would ask someone to remove the obstruction. Can MLLMs exhibit a similar “proactive” behavior by requesting simple user interventions? To investigate this, we introduce ProactiveBench, a benchmark built from seven repurposed… See the full description on the dataset page: https://huggingface.co/datasets/tdemin16/ProactiveBench.Proactive-CSI-ProcessedROMA_proactive
ROMA Proactive Streaming Dataset
Figure: Overview of ROMA's Streaming Dataset. This repository contains the Proactive subset (Green and Purple sections).
Dataset Summary
This repository contains the Proactive Interaction subset of the dataset introduced in the paper ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding.
This dataset is designed to train multimodal models for streaming video understanding, specifically focusing on tasks… See the full description on the dataset page: https://huggingface.co/datasets/EurekaTian/ROMA_proactive.trainProactiveMobile
ProactiveMobile
A comprehensive, executable benchmark for proactive intelligence in mobile agents — agents that anticipate user needs and act on their own, rather than passively executing explicit commands.
📄 Paper: arXiv:2602.21858 · 🔗 Project: xiaomi-research/proactive-mobile
Overview
Each instance asks a model to infer latent user intent from four dimensions of on-device context, then produce an executable function sequence drawn from a unified function… See the full description on the dataset page: https://huggingface.co/datasets/xiaomi-research/ProactiveMobile.
