datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Vision2Web
Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
[🏠 Project Page] [📖 arXiv Paper] [🏆 Leaderboard] [📮 Submit Results]
Vision2Web is a comprehensive benchmark designed to evaluate multimodal coding agents on visual website development tasks spanning the full software development lifecycle.
This dataset repository contains the benchmark tasks, UI prototypes, test workflows, and resources used to evaluate agent performance.… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/Vision2Web.SurveyReview
SurveyReview
SurveyReview is a reviewer-aligned benchmark for evaluating survey papers. It turns real peer-review reports into multidimensional scores and rationales so that model judgments can be compared with human reviewer judgments.
The benchmark covers four dimensions: Readability, Criticalness, Comprehensiveness, and Structure.
Latest release: v1.1
What's New in v1.1
Cleaned full-text content for 1,646 survey articles, stored in two JSON shards.
The… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/SurveyReview.
