Lexmount/LexBench-Browser
LexBench-Browser LexBench-Browser is a public browser-agent dataset for evaluating agents on real-web workflows. The v1.0 snapshot contains 210 no-login tasks across 107 distinct websites, with Chinese and English instructions, task-level reference steps, key points, common mistakes, scoring rubrics, and robustness tags. Repository: https://github.com/lexmount/browseruse-agent-bench Dataset page: https://huggingface.co/datasets/Lexmount/LexBench-Browser Docs:… See the full description on the dataset page: https://huggingface.co/datasets/Lexmount/LexBench-Browser.
Update dataset card
Update LexBench-Browser dataset files
Fix iQiyi target website in LexBench Browser
Fix iQiyi target website in LexBench Browser
Restore China politics tasks in public LexBench-Browser data
Fix dataset loading config after task removal
Fix LexBench-Browser public dataset loading config
Update LexBench-Browser public dataset card after task removal
Remove China politics tasks from LexBench-Browser public data
Update LexBench-Browser dataset files
Update LexBench-Browser dataset files
Update dataset card
Clarify dataset card naming
Fix dataset card metadata
Update dataset card
Update LexBench-Browser dataset files
docs: bump README to v3.1 stats (210 tasks; document risk_control / multi-site target_website)
data: L1 quality fix — port PR #226 (210 records; scoring/metadata/risk_control updates; 7 records moved out)
docs: update README counts to reflect 217 task release
data: drop login_required tasks (220 -> 217)
docs: clean README — drop references to non-public mirror
docs: refresh README for v3.0 (single task.jsonl, 220 tasks, no scenario_tier)
v3.0: consolidate to task.jsonl (L1+L3, 220 tasks); drop scenario_tier field
Restore L1 dataset robustness tags.
Upload l1.jsonl
Upload folder using huggingface_hub
initial commit
