datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Web-Bench
Web-Bench
English | 中文 README
📖 Overview
Web-Bench is a benchmark designed to evaluate the performance of LLMs in actual Web development. Web-Bench contains 50 projects, each consisting of 20 tasks with sequential dependencies. The tasks implement project features in sequence, simulating real-world human development workflows. When designing Web-Bench, we aim to cover the foundational elements of Web development: Web Standards and Web Frameworks. Given the scale and… See the full description on the dataset page: https://huggingface.co/datasets/bytedance-research/Web-Bench.web-research-trajectories
Web-Research Agent Trajectories
The first open dataset from Assayo — an open rubric and method for judging the quality
of AI agent trajectories. (The name is from assay*: to test the purity of a metal.)*
An open rubric and a small, hand-built gold set for judging multi-step web-research
agent trajectories. A trajectory is the full record of an agent solving one task by
searching the web, reading sources, and answering with citations — the
think → act → observe → repeat →… See the full description on the dataset page: https://huggingface.co/datasets/Assayo/web-research-trajectories.
