datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cua-bench
CUA-Bench
CUA-Bench contains 90 browser-agent evaluation tasks selected from
Online-Mind2Web.
The original task IDs, instructions, and initial website URLs are preserved.
Splits
Split
Task type
Rows
navigation
Navigate to a requested page or website state
30
single_search
Complete one direct search or lookup
30
multi_step_search
Complete a search requiring multiple steps or constraints
30
Total
90
Row schema
Field… See the full description on the dataset page: https://huggingface.co/datasets/WootzappLab/cua-bench.cua-bench
CUA-Bench: 75 original CUA-Gym tasks
A selection of 75 unchanged original tasks from xlangai/CUA-Gym, grouped into three evaluation splits of 25 tasks each. Every task includes its original instruction, setup script and matching reward.py. No synthetic task variants or replacement grading rules were created.
This repository also includes the source code for the 13 required mock websites, a Docker environment, and a controller helper for running original setup and reward code.… See the full description on the dataset page: https://huggingface.co/datasets/ishagarg1103/cua-bench.
