ProCreations/betterwright-agentic-browser-50k
BetterWright Agentic Browser — 6,093-row stopped checkpoint This is the public checkpoint of a generation run originally planned for 50,000 rows. Generation was stopped at the account owner's request and the exact 6,093 accepted rows were packaged. It is synthetic training data, not live browser recordings. Contents 5,971 BetterWright demonstrations and 122 Playwright demonstrations. 32 task domains and 21 browser feature categories. Harness-shaped conversations… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/betterwright-agentic-browser-50k.
BetterWright Agentic Browser — 6,093-row stopped checkpoint
This is the public checkpoint of a generation run originally planned for 50,000 rows. Generation was stopped at the account owner's request and the exact 6,093 accepted rows were packaged. It is synthetic training data, not live browser recordings.
Contents
- 5,971 BetterWright demonstrations and 122 Playwright demonstrations.
- 32 task domains and 21 browser feature categories.
- Harness-shaped conversations for Codex CLI, OpenCode, Pi, and Hermes.
- Compact
OBS; NEXT; CHECK; FALLBACKdecision rationale. - Train/validation/test split: 5,851/121/121.
- Generator and critic:
qwen3.8-27b. Retained rows scored at least 8/10 per automated category and averaged at least 8.6/10.
Each row includes id, task metadata, user_request, start_url, a uniform messages list, automated quality_scores, quality_issues, and generation metadata. Tool arguments are JSON strings.
Important limitations
The trajectories and web states are simulated on reserved .example hosts; they were not replayed against live sites. The generator and critic used the same model, so the scores are not an independent evaluation. A checkpoint audit found that many concise simulated tool results do not expose full BetterWright UI/control directories, and some later actions use CSS/test-id selectors not explicitly present in prior observations. This can teach selector hallucination. Treat the corpus as experimental synthetic SFT data, mix it with execution-verified trajectories, and evaluate downstream agents in a sandbox before consequential browser use.
Validation
The packaged checkpoint has 6,093 unique IDs and 6,093 unique normalized user requests. Message/tool-call pairing, JSON tool results, reasoning format, score thresholds, and split row counts were checked before upload. See validation_report.json and generation_config.json.
Source grounding
Tool and harness shapes were developed against pinned public source revisions:
{
"betterwright": "5ba586ff1b26353774d81d6b8f8241d0969251c5",
"codex": "eb10d91e48ccbd0930427461fb392337addb1ac0",
"hermes-agent": "df4b3733baa5534bec1e6e888e5ba86a0a6f9c3b",
"opencode": "69c172e8a7c0086887b1f93ed5a162f14b6aa0c5",
"pi-mono": "96317e50b8d6e7f6d0e47fd29122baf1461c00f5"
}BetterWright is an MIT-licensed project from The BetterWright Project and contributors. This independently generated dataset is not an official BetterWright release: https://github.com/BetterWright/betterwright
