datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Vision2Web
Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
[🏠 Project Page] [📖 arXiv Paper] [🏆 Leaderboard] [📮 Submit Results]
Vision2Web is a comprehensive benchmark designed to evaluate multimodal coding agents on visual website development tasks spanning the full software development lifecycle.
This dataset repository contains the benchmark tasks, UI prototypes, test workflows, and resources used to evaluate agent performance.… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/Vision2Web.CC-Bench-trajectories
CC-Bench Trajectories Overview
To evaluate GLM-4.6's agentic coding capabilities in real-world scenarios, we developed CC-Bench-V1.1 using Claude Code as the agentic coding testbed. Building on CC-Bench-V1.0, we added 22 more challenging coding tasks and conducted comprehensive evaluations against Claude-Sonnet-4, GLM-4.5, Kimi-K2-0905, and DeepSeek-V3.1-Terminus. The benchmark comprises 74 coding tasks spanning frontend development, tool development, data analysis, testing, and… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/CC-Bench-trajectories.RuHeritage-Corpus
RuHeritage-Corpus
🇬🇧 English Description
RuHeritage-Corpus is a high-quality, curated dataset of Russian classical literature, specifically designed for the pre-training and continued pre-training (CPT) of Large Language Models (LLMs).
The corpus focuses on the Golden and Silver Ages of Russian literature, providing models with exposure to rich vocabulary, complex syntactic structures, and stylistically flawless Russian text, acting as a "quality anchor"… See the full description on the dataset page: https://huggingface.co/datasets/zait-ai/RuHeritage-Corpus.OpenJA
OpenJA
🇬🇧 English Description
OpenJA is a dataset of clean, officially published parliamentary transcripts of the Japanese language. It is characterized by high-quality text (without web noise, HTML, advertising) and reliable metadata, but it represents one narrow language register (official/parliamentary speech) and is better suited as an addition to more diverse corpora than as the only source for a general-purpose pretrain.
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/zait-ai/OpenJA.
