datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
browser-agent-tasks
Browser Agent Tasks
Moderate, multi-step browser tasks for collecting agent trajectories and evaluating screenshot, DOM, and DOM-diff evidence.
Files
tasks.jsonl: one task definition per line.
Each task contains:
task_id: stable identifier
category: task family
instruction: complete instruction given to the browser agent
start_url: suggested public starting page
stopping_condition: when the agent must stop
constraints: safety and scope restrictions
These tasks… See the full description on the dataset page: https://huggingface.co/datasets/ishagarg1103/browser-agent-tasks.browser-agent-failure-corpus
Browser Agent Failure Corpus — case-row view
This view contains 30 controlled local regression cases from one owner-produced deterministic run on 2026-08-31 using Qwen/Qwen3-8B-GGUF:Q4_K_M, llama.cpp b10103-c588c4f47, the Codex in-app browser and the historical Mutant Web guard v2, at temperature 0, reasoning budget 0 and at most eight steps per case. It preserves model proposals, guard decisions, executed actions and recorded outcomes for inspectability and reuse. Allowed… See the full description on the dataset page: https://huggingface.co/datasets/mutantweb/browser-agent-failure-corpus.browser-agent-tasks
Browser Agent Tasks
Moderate, multi-step browser tasks for collecting agent trajectories and evaluating screenshot, DOM, and DOM-diff evidence.
Files
tasks.jsonl: one task definition per line.
Each task contains:
task_id: stable identifier
category: task family
instruction: complete instruction given to the browser agent
start_url: suggested public starting page
stopping_condition: when the agent must stop
constraints: safety and scope restrictions
These tasks… See the full description on the dataset page: https://huggingface.co/datasets/WootzappLab/browser-agent-tasks.Agent-browser-task
