datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-Dev-train📝 Paper | 🌐 Github
🤗 SWE-Dev-7B (Qwen-2.5-Coder-7B-Instruct)
🤗 SWE-Dev-9B (GLM-4-9B-Chat)
🤗 SWE-Dev-32B (Qwen-2.5-Coder-32B-Instruct)
🤗 SWE-Dev-train (Training Data)
🚀 SWE-Dev, an open-source Agent for Software Engineering tasks! This repository contains the SWE-Dev-32B model as presented in the paper SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling.
💡 We develop a comprehensive pipeline for creating developer-oriented datasets from GitHub… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/SWE-Dev-train.Vision2Web
Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
[🏠 Project Page] [📖 arXiv Paper] [🏆 Leaderboard] [📮 Submit Results]
Vision2Web is a comprehensive benchmark designed to evaluate multimodal coding agents on visual website development tasks spanning the full software development lifecycle.
This dataset repository contains the benchmark tasks, UI prototypes, test workflows, and resources used to evaluate agent performance.… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/Vision2Web.fineweb2-arb-eduglm-simple-evals-dataset
glm-simple-evals-dataset
This repository is dedicated to storing various evaluation data required for the glm-simple-evals evaluation project, to enable industry researchers and developers to reproduce the performance of the GLM-4.5 series models on reported benchmarks.
Currently, this repository covers the data required for the following evaluation tasks:
AIME
GPQA
HLE
LiveCodeBench
MATH 500
SciCode
MMLU Pro
Usage Instructions
To use these evaluation datasets… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/glm-simple-evals-dataset.CC-Bench-trajectories
CC-Bench Trajectories Overview
To evaluate GLM-4.6's agentic coding capabilities in real-world scenarios, we developed CC-Bench-V1.1 using Claude Code as the agentic coding testbed. Building on CC-Bench-V1.0, we added 22 more challenging coding tasks and conducted comprehensive evaluations against Claude-Sonnet-4, GLM-4.5, Kimi-K2-0905, and DeepSeek-V3.1-Terminus. The benchmark comprises 74 coding tasks spanning frontend development, tool development, data analysis, testing, and… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/CC-Bench-trajectories.RuHeritage-Corpus
RuHeritage-Corpus
🇬🇧 English Description
RuHeritage-Corpus is a high-quality, curated dataset of Russian classical literature, specifically designed for the pre-training and continued pre-training (CPT) of Large Language Models (LLMs).
The corpus focuses on the Golden and Silver Ages of Russian literature, providing models with exposure to rich vocabulary, complex syntactic structures, and stylistically flawless Russian text, acting as a "quality anchor"… See the full description on the dataset page: https://huggingface.co/datasets/zait-ai/RuHeritage-Corpus.AISE-Bench
AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs
🌐 Project Page •
💻 GitHub •
📖 KDD 2026 Paper
AISE-Bench is a real-world benchmark for information seeking on academic knowledge graphs. It is built from authentic AMiner user search queries and provides human-verified academic question-answering data with executable multi-step API trajectories, standardized tool inputs, API execution outputs, and… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/AISE-Bench.IKMKLThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 10,
"total_frames": 13562,
"total_tasks":1,
"total_videos": 30,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zaidyzahar/IKMKL.OpenJA
OpenJA
🇬🇧 English Description
OpenJA is a dataset of clean, officially published parliamentary transcripts of the Japanese language. It is characterized by high-quality text (without web noise, HTML, advertising) and reliable metadata, but it represents one narrow language register (official/parliamentary speech) and is better suited as an addition to more diverse corpora than as the only source for a general-purpose pretrain.
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/zait-ai/OpenJA.RusWeather-434
RusWeather-434
🇬🇧 English
Dataset Summary
RusWeather-434 is a daily weather time-series dataset covering 434 cities
across Russia, from 2019-01-01 to the 2026.07.24. For each city and each
day, the dataset provides temperature, humidity, precipitation, wind, surface
pressure, solar radiation, and cloud cover.
Total Records: 1,197,406.
Total size: 23.8 MB.
Format: Parquet.
Data is sourced from NASA POWER (Prediction Of Worldwide Energy
Resources)… See the full description on the dataset page: https://huggingface.co/datasets/zait-ai/RusWeather-434.so100_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 1600,
"total_tasks":1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zainhub/so100_test.CIS-Weather-807
CIS-Weather-807
Russia · 53.98% · 1,197,406 rows
Azerbaijan · 6.47% · 143,468 rows
Uzbekistan · 6.34% · 140,709 rows
Belarus · 6.34% · 140,709 rows
Kazakhstan · 6.22% · 137,950 rows
Moldova · 5.60% · 124,155 rows
Kyrgyzstan · 5.10% · 113,119 rows
Armenia · 4.98% · 110,360 rows
Tajikistan · 4.98% · 110,360 rows
Total · 2,218,236 rows… See the full description on the dataset page: https://huggingface.co/datasets/zait-ai/CIS-Weather-807.BigEarthNet.txt
BigEarthNet.txt: A Large-Scale Multi-Sensor Image-Text Dataset and Benchmark for Earth Observation
BigEarthNet.txt is a large-scale multi-sensor image–text dataset for Earth observation, designed to advance vision–language learning on remote sensing data. It comprises 464,044 co-registered Sentinel-1 (SAR) and… See the full description on the dataset page: https://huggingface.co/datasets/zaidxx8/BigEarthNet.txt.java-vulnerabilityworld-models-eval
DreamGrasp: Processed LIBERO Manipulation Demonstrations
Does a robot policy's evaluation still mean something if it never touched a real simulator, only a world model's imagination of one?
This dataset is the shared training data behind that question, a single, ready-to-train release built from LIBERO's manipulation demonstrations (libero_spatial, libero_object, libero_goal). It provides:
Fixed, versioned train / validation / test / held-out splits, so every result trained on… See the full description on the dataset page: https://huggingface.co/datasets/ZaidGhazal/world-models-eval.Duolingo-Spaced-Repetition-Dataexpresso-tagged-w-speech-gemmazai-org_SWE-Dev-train_formattedCGSQuAD
Dataset Card for "CGSQuAD"
More Information needed
osworld_tasks_filesBPtelecom-churn-datasettau2-bench_zai-org_GLM-4.5-Air_n1_r1WeatherHumanoidRLdqtraserz-ai-glm-4.5-air-nt3-T1-H100expresso-tagsexpresso-finalexpresso-tags-with-defaults
