web-agents
webagentsexecution-time-warnings-web-agents
Trustworthy Completion for Web Agents
This release contains audited run-level results from a controlled study of execution-time safeguards for web agents under deceptive consumer interfaces.
The benchmark independently scores nominal completion (C) and trajectory safety (S): trustworthy completion (C=1,S=1), unsafe completion (C=1,S=0), safe non-completion (C=0,S=1), and unsafe failure (C=0,S=0).
Study design
One frozen vision-capable web-agent configuration
12… See the full description on the dataset page: https://huggingface.co/datasets/deceptive-web-benchmark/execution-time-warnings-web-agents.Execution-Time-Warnings-for-Web-Agents
Trustworthy Completion for Web Agents
This release contains audited run-level results from a controlled study of execution-time safeguards for web agents under deceptive consumer interfaces.
The benchmark independently scores nominal completion (C) and trajectory safety (S): trustworthy completion (C=1,S=1), unsafe completion (C=1,S=0), safe non-completion (C=0,S=1), and unsafe failure (C=0,S=0).
Study design
One frozen vision-capable web-agent configuration
12… See the full description on the dataset page: https://huggingface.co/datasets/EvaNing123/Execution-Time-Warnings-for-Web-Agents.execution-time-warnings-web-agents
Deception Warning Study — benchmark runs (staging)
This repository will host run-level rows for the controlled benchmark described in the companion paper (NeurIPS-style release).
Contents (when populated)
Artifact
Description
run_level.jsonl / run_level.csv
One row per merged run: task, condition, repeat, outcome, flags
run_level.parquet
Optional if pyarrow is installed (Hub-friendly)
manifest.yaml (optional)
benchmark_version, repeats_per_task_condition… See the full description on the dataset page: https://huggingface.co/datasets/deceptive-web/execution-time-warnings-web-agents.web_agents_google_flight_trajectories
Web Agent Google Flight Trajectories
This dataset was originally created on Nov 23 2024 during EF's Ai On Edge Hackathon.
The purpose of this dataset is to give both positive and negative image web-agent trajectories to finetune small edge-models on web agentic tasks.
Web-Agent-SearXNG
