datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
webvoyager_evaluation_dataWebVoyager-Trajectories-GPT-4V
The trajectories of WebVoyager (GPT-4V as backbone)
Paper: WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
Code: https://github.com/MinorJerry/WebVoyager
WebVoyager2025Valid
WebVoyager 2025 Valid
A modified subset of WebVoyager designed to be valid until 20th December 2025.
This was used to benchmark proxy-lite.
You can find the original WebVoyager tasks here.
webvoyagerweb-voyagerWebVoyager-GAIA-SparseAttention-Offline-Qwen3-VL-30B-A3B
Sparse Attention for Web Agents — Offline Replay Records
Step-level records from an offline replay study comparing three sparse-attention methods against
full attention on browser-agent trajectories.
Model: Qwen3-VL-30B-A3B-Instruct (48 layers, 128 experts / 8 active, GQA 32:4, head_dim 128)
Hardware: NVIDIA GB10 (sm_121)
Tasks: 50 complete trajectories sampled from WebVoyager + GAIA — 332 steps, of which 190 emit an
element index. Every task was completed successfully by the… See the full description on the dataset page: https://huggingface.co/datasets/shiqihe/WebVoyager-GAIA-SparseAttention-Offline-Qwen3-VL-30B-A3B.WebVoyager-Trajectories-Qwen3.5-Omni
WebVoyager + GAIA Agent Trajectories (Qwen3.5-Omni)
Full browser-agent trajectories for 733 tasks (643 WebVoyager + 90 GAIA-web), produced by a browser-use agent driven by qwen3.5-omni-plus-2026-03-15 (multimodal, via Alibaba DashScope), run in a real headed browser. Every step records the exact LLM context (including the screenshot the model saw) and the action taken, plus a reference-grounded success verdict.
Results
Judged by the WebVoyager reference-grounded… See the full description on the dataset page: https://huggingface.co/datasets/shiqihe/WebVoyager-Trajectories-Qwen3.5-Omni.
