datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
st-webagentbench
A Benchmark for Evaluating Safety & Trustworthiness in Web Agents
Accepted at ICLR 2026
Overview
ST-WebAgentBench is a policy-enriched evaluation suite for web agents, built on BrowserGym. It measures not only whether agents complete tasks, but whether they do so while respecting safety and trustworthiness (ST) policies — the constraints that govern real enterprise deployments.
The… See the full description on the dataset page: https://huggingface.co/datasets/ST-WebAgentBench/st-webagentbench.agent-web-index
Agent Web Index — how much of the web can AI assistants actually read?
48,154 domains measured live. 25% of them cannot be read by at least one of
ChatGPT, Claude, Perplexity or Gemini. Updated daily. Live index: https://shop.lumnika.com/ai-readiness/
Every row here is the result of real HTTP requests, not an estimate and not a re-publication of
someone else's crawl: each domain's homepage is requested once as a browser and once as each of the
published AI crawler user-agents… See the full description on the dataset page: https://huggingface.co/datasets/DeusHorizon/agent-web-index.embodied-web-agent-geoguessr
