datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
browser-tool-use-trajectories
Browser Tool Use Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/browser-tool-use-trajectories.BrowserAgent-Data
BrowserAgent ChatML Dataset (SFT/RFT)
This dataset contains ChatML-style multi-turn dialogues for a browser agent task. The data is prepared as JSON Lines so it can be previewed directly with the Hugging Face Hub Data Visualizer and loaded with the datasets library.
Links
Paper
Github
Files
sft.jsonl — SFT split (one JSON object per line)
rft.jsonl — RFT split (one JSON object per line)
Schema
Each record is a JSON object containing:
messages:… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/BrowserAgent-Data.browser-agent-tasks
Browser Agent Tasks
Moderate, multi-step browser tasks for collecting agent trajectories and evaluating screenshot, DOM, and DOM-diff evidence.
Files
tasks.jsonl: one task definition per line.
Each task contains:
task_id: stable identifier
category: task family
instruction: complete instruction given to the browser agent
start_url: suggested public starting page
stopping_condition: when the agent must stop
constraints: safety and scope restrictions
These tasks… See the full description on the dataset page: https://huggingface.co/datasets/ishagarg1103/browser-agent-tasks.browser-agent-failure-corpus
Browser Agent Failure Corpus — case-row view
This view contains 30 controlled local regression cases from one owner-produced deterministic run on 2026-08-31 using Qwen/Qwen3-8B-GGUF:Q4_K_M, llama.cpp b10103-c588c4f47, the Codex in-app browser and the historical Mutant Web guard v2, at temperature 0, reasoning budget 0 and at most eight steps per case. It preserves model proposals, guard decisions, executed actions and recorded outcomes for inspectability and reuse. Allowed… See the full description on the dataset page: https://huggingface.co/datasets/mutantweb/browser-agent-failure-corpus.browser-agent-tasks
Browser Agent Tasks
Moderate, multi-step browser tasks for collecting agent trajectories and evaluating screenshot, DOM, and DOM-diff evidence.
Files
tasks.jsonl: one task definition per line.
Each task contains:
task_id: stable identifier
category: task family
instruction: complete instruction given to the browser agent
start_url: suggested public starting page
stopping_condition: when the agent must stop
constraints: safety and scope restrictions
These tasks… See the full description on the dataset page: https://huggingface.co/datasets/WootzappLab/browser-agent-tasks.betterwright-agentic-browser-50k
BetterWright Agentic Browser — 6,093-row stopped checkpoint
This is the public checkpoint of a generation run originally planned for 50,000 rows. Generation was stopped at the account owner's request and the exact 6,093 accepted rows were packaged. It is synthetic training data, not live browser recordings.
Contents
5,971 BetterWright demonstrations and 122 Playwright demonstrations.
32 task domains and 21 browser feature categories.
Harness-shaped conversations… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/betterwright-agentic-browser-50k.browserbench-live
BrowserBench Live Tasks
This private dataset contains 292 live-browser task-generation records derived only from the starting_url field of Halluminate/BrowserBench at revision aa56ce5e6331425c29878037fa8c169507deecdc.
Pages were visited live at generation time. Historical prompts, results, and ground-truth URLs were not used. Every source row is retained, including rejected pages, with its rejection reason. Fresh page screenshots are stored under screenshots/.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/merve/browserbench-live.awesome-browser-use-prompts
Awesome Browser-Use Prompts
A curated collection of effective prompts for Browser-Use, the framework that enables AI agents to control web browsers. This repository aims to provide examples, templates, and best practices for crafting prompts that maximize the capabilities of Browser-Use agents.
Introduction
Browser-Use allows language models to interact with web interfaces through natural language instructions. Effective prompting is crucial to achieving successful… See the full description on the dataset page: https://huggingface.co/datasets/metehan777/awesome-browser-use-prompts.tool-calling-browser-agent-tasks
Dataset Card
Created by: DataCreator AI
Overview
Tool Calling for Agentic Tasks with Multi-Step Workflows contains 1,062 synthetic multi-turn conversations between a user and an AI assistant. The examples primarily focus on practical agentic tasks such as train ticket booking, dynamic form filling, and payment processing. It provides diverse scenarios including successful execution, context retrieval, tool integration, and failure recovery.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/DataCreatorAI/tool-calling-browser-agent-tasks.browserbench-live-expanded
BrowserBench Live Expanded
Live-browser tasks derived only from the starting_url field of
Halluminate/BrowserBench at revision
aa56ce5e6331425c29878037fa8c169507deecdc.
Each source page was revisited live at 800×600. Up to three tasks were generated
from the current page observation: information, navigation, and interaction.
Historical prompts, results, screenshots, and ground-truth URLs were not used.
The canonical dataset contains all 876 task slots from 292 source pages. A… See the full description on the dataset page: https://huggingface.co/datasets/merve/browserbench-live-expanded.2d-webmcp-browser-focus
2D WebMCP Browser Focus (Prerelease)
What this is
This is an early test of whether agents need useful tool results to complete an accessible browser task.
The agent must add a Retry step to a workflow, connect it correctly, and move keyboard focus to that new step. The test checks the real browser, not just the agent's final answer.
What happened
We ran each version 20 times with gpt-5-mini using low reasoning effort.
Tool result
Verified… See the full description on the dataset page: https://huggingface.co/datasets/accesslint/2d-webmcp-browser-focus.agentic-tooluse-computer-browserbrowseragent-dataAgent-browser-taskminicpm5-computer-browser-coding-v2probe-browser-evaltiny-browser-planner-reason-dataset
language:
en
task_categories:
text-generation
tags:
reasoning
planning
browser-agent
build-small
size_categories:
n<1K
pretty_name: TinyBrowserPlanner Reason Dataset
TinyBrowserPlanner-Reason Dataset
Dataset used to train the Reason-First version of TinyBrowserPlanner.
Key Result
4/12 → 10/12 planning accuracy
Contains
reasoning examples
replanning scenarios
wrong-page recovery
paywall recovery
refine_search examples… See the full description on the dataset page: https://huggingface.co/datasets/Georgefifth/tiny-browser-planner-reason-dataset.
