datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mind2Web
Dataset Card for Dataset Name
Dataset Summary
Mind2Web is a dataset for developing and evaluating generalist agents for the web that can follow language instructions to complete complex tasks on any website. Existing datasets for web agents either use simulated websites or only cover a limited set of websites and tasks, thus not suitable for generalist web agents. With over 2,000 open-ended tasks collected from 137 websites spanning 31 domains and crowdsourced action… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/Mind2Web.Online-Mind2Web
Blog |
Paper |
Code |
Leaderboard
Online-Mind2Web
Online-Mind2Web is the online version of Mind2Web, a more diverse and user-centric dataset includes 300 high-quality tasks from 136 popular websites across various domains. The dataset covers a diverse set of user tasks, such as clothing, food, housing, and transportation, to evaluate web agents' performance in a real-world online environment.
News
[11/03/2025] We’ve updated 36 tasks that are… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/Online-Mind2Web.MT-Mind2Web
MT-Mind2Web Dataset
MT-Mind2Web is constructed by using the single-turn interactions from Mind2Web, an expert-annotated web navigation dataset, as the guidance to construct conversation sessions.
Statistics
Train
Test-Task
Test-Website
Test-Subdomain
# Conversations
600
34
42
44
# Turns
2,896
191
218
216
Avg. # Turn/Conv.
4.83
5.62
5.19
4.91
Avg. # Action/Turn
2.95
3.16
3.01
3.07
Avg. # Element/Turn
573.8
626.3
620.6
759.4
Avg. Inst. Length
36.3… See the full description on the dataset page: https://huggingface.co/datasets/magicgh/MT-Mind2Web.Mind2Web-LiveDataset: Mind2Web-Live
Github: https://github.com/iMeanAI/WebCanvas
Mind2Web_train_llava
Mind2Web training set for the paper: Harnessing Webpage Uis For Text Rich Visual Understanding
🌐 Homepage | 🐍 GitHub | 📖 arXiv
Introduction
We introduce MultiUI, a dataset containing 7.3 million samples from 1 million websites, covering diverse multi- modal tasks and UI layouts. Models trained on MultiUI not only excel in web UI tasks—achieving up to a 48% improvement on VisualWebBench and a 19.1% boost in action accuracy on a web agent dataset Mind2Web—but also… See the full description on the dataset page: https://huggingface.co/datasets/neulab/Mind2Web_train_llava.JevForge-Mind2Web
JevForge Mind2Web Gold Decisions
Private research snapshot of gold candidate-decision records used by
JevForge.
What is included
Each JSONL record contains a page state, a choice question over candidate
elements, a noul question about one candidate, complete gold target
distributions, a website group, and source annotation metadata.
Split
Records
Websites
train
4,642
49
dev
786
6
calibration
400
6
test
800
8
ood
386
4
The 7,014 record IDs… See the full description on the dataset page: https://huggingface.co/datasets/AndeyTait/JevForge-Mind2Web.i-Mind2Webnull
Sequence-of-action-prediction-mind2webmind2web_train_configsMind2Web_AXTmind2web_evaluation_dataMind2Web_with_act_descWe added an "action_description" field to each action, generated by GPT-4, which provides a semantic description of the action.
Semantic Annotation Construction: Following the approach used in the second stage of the AITZ dataset, we designed multimodal inputs for each operation in the Mind2Web dataset. These inputs consist of two screenshots and a prompt:
The first screenshot is the original webpage interface.
The second screenshot shows the original interface with a blue plus sign marking… See the full description on the dataset page: https://huggingface.co/datasets/xzq11111/Mind2Web_with_act_desc.mind2web-sharegpt
