datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MMRC_Real_World_Conversation
MMRC - Multi-Modal Open-Ended Conversation Dataset
Overview:
MMRC is a benchmark dataset designed for evaluating Multi-Modal Large Language Models (MLLMs) in open-ended, multi-turn conversations. It provides diverse, real-world conversational data that integrates both textual and visual modalities, aiming to push the boundaries of MLLM performance in practical settings.
Dataset Details:
The MMRC dataset is composed of multi-turn conversations with integrated… See the full description on the dataset page: https://huggingface.co/datasets/WUUE/MMRC_Real_World_Conversation.bmvs_sparse_dtu2VeriOS-Bench
VeriOS-Bench
Paper | Code
Dataset Overview
This dataset is a Query-Driven Trustworthy OS Agent Dataset implemented as described in the paper:
VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents
The VeriOS dataset is designed to train and evaluate operating system (OS) agents that automate tasks through on-device graphical user interfaces (GUIs). It focuses on enabling agents to decide when to query humans for more reliable task completion… See the full description on the dataset page: https://huggingface.co/datasets/wuuuuuz/VeriOS-Bench.ESPPMobileIAR
Dataset Overview
This dataset is a User-Specific Personalized GUI Mobile Agents Dataset implemented as described in the paper:
Quick on the Uptake: Eliciting Implicit Intents from Human Demonstrations for Personalized Mobile-Use Agents
Citation
@article{wu2025quick,
title={Quick on the Uptake: Eliciting Implicit Intents from Human Demonstrations for Personalized Mobile-Use Agents},
author={Wu, Zheng and Huang, Heyuan and Yang, Yanjia and Song, Yuanyi and Lou, Xingyu… See the full description on the dataset page: https://huggingface.co/datasets/wuuuuuz/MobileIAR.highresdtutest
