datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LongBench-v2
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
🌐 Project Page: https://longbench2.github.io
💻 Github Repo: https://github.com/THUDM/LongBench
📚 Arxiv Paper: https://arxiv.org/abs/2412.15204
LongBench v2 is designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world multitasks. LongBench v2 has the following features: (1) Length: Context length ranging from 8k to… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongBench-v2.LongBenchLongBench is a comprehensive benchmark for multilingual and multi-task purposes, with the goal to fully measure and evaluate the ability of pre-trained language models to understand long text. This dataset consists of twenty different tasks, covering key long-text application scenarios such as multi-document QA, single-document QA, summarization, few-shot learning, synthetic tasks, and code completion.terminal-bench-2-verified
Terminal-Bench 2.0 Verified: Instruction & Environment Fix Version
中文版本
2026.08.18 Update: Following the release of Terminal-Bench 2.1, we conducted another round of review on several tasks' grading and instructions, validating each fix end-to-end inside the actual task images. This round fixes 6 tasks in two categories: Grading/Test Fixes (dna-insert, make-doom-for-mips, filter-js-from-html, install-windows-3.11) — the test logic itself misjudged valid submissions or let… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/terminal-bench-2-verified.Quranic-Recitation-Data
🌟 Overview
Quranic Recitation Dataset (Word-by-Word Sync) is a highly optimized, production-ready dataset containing high-quality audio recitations of the Holy Quran synchronized at the word-by-word level.
This dataset features 135 world-renowned reciters, with every Surah (114 chapters) mapped precisely to millisecond-accurate word timestamps. It is designed for modern Islamic mobile and web applications — served via a Cloudflare Edge CDN with native… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Recitation-Data.MotionBench
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
[🍎 Project Page] [📖 arXiv Paper] [📊 Dataset] [💻 GitHub] [🏆 Leaderboard] [🏆 HF Leaderboard]
MotionBench is a comprehensive evaluation benchmark designed to assess the fine-grained motion comprehension of video understanding models. It evaluates models' motion-level perception through six primary categories of motion-oriented question types and includes data… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/MotionBench.LongAlign-10k
LongAlign-10k
🤗 [LongAlign Dataset] • 💻 [Github Repo] • 📃 [LongAlign Paper]
LongAlign is the first full recipe for LLM alignment on long context. We propose the LongAlign-10k dataset, containing 10,000 long instruction data of 8k-64k in length. We investigate on trianing strategies, namely packing (with loss weighting) and sorted batching, which are all implemented in our code. For real-world long context evaluation, we introduce LongBench-Chat that evaluate the… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongAlign-10k.Quranic-Translation-Audio-Data
Overview
Quranic Translation Audio Data is a highly curated, standardized, and streaming-optimized multilingual audio dataset containing the complete recitation of translation audios and commentaries of the Holy Quran across 51 different translation directories.
Every audio track has been meticulously converted from heavy .mp3 source files into the modern, high-fidelity Opus (.opus) format at a streaming-optimized bitrate of 32kbps. Alongside… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Translation-Audio-Data.Quranic-Word-By-Word-Audio-Data
🌟 Overview
Quran Word-By-Word Audio Dataset contains two complete word-by-word recitation datasets of the Holy Quran, optimized for edge delivery, mobile streaming, and machine learning pipelines:
Muallim (Teacher Style) — optimized for slow, educational, and repeat-friendly listening.
Mujawwad (Tajweed Style) — optimized for natural rhythmic recitation with full tajweed flow.
Originally averaging between 2.0 GB to 2.3 GB each in raw format, the… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Word-By-Word-Audio-Data.DeepDive
DeepDive Dataset
Overview
This is the training dataset for DeepDive, an automated approach for training deep search agents with complex, multi-step reasoning capabilities. The dataset is constructed through automated knowledge graph random walks, entity obfuscation, and difficulty filtering to create challenging questions that require sophisticated search and retrieval skills.
Dataset Statistics
Component
Split
Size
Description
Total… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/DeepDive.RPC-Bench
RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension
🌐 Project Page •
💻 GitHub •
📖 Paper
RPC-Bench is a fine-grained benchmark for research paper comprehension. It is built from review-rebuttal exchanges of high-quality academic papers and supports both text-only and visual evaluation through complementary paper representations.
Data Structure
RPC-Bench is organized into train, dev, and test subsets. Split assignments… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/RPC-Bench.mmlu-random-Ahumaneval-xHumanEval-X is a benchmark for the evaluation of the multilingual ability of code generative models. It consists of 820 high-quality human-crafted data samples (each with test cases) in Python, C++, Java, JavaScript, and Go, and can be used for various tasks.AgentInstruct
AgentInstruct Dataset
🤗 [Models] • 💻 [Github Repo] • 📌 [Project Page] • 📃 [Paper]
AgentInstruct is a meticulously curated dataset featuring 1,866 high-quality interactions, designed to enhance AI agents across six diverse real-world tasks, leveraging innovative methods like Task Derivation and Self-Instruct.
🔍 CoT - Harness the power of ReAct, offering detailed thought explanations for each action, ensuring an intricate understanding of the model's decision-making… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/AgentInstruct.Quranic-Recitation-AlignmentImageRewardDBImageRewardDB is a comprehensive text-to-image comparison dataset, focusing on text-to-image human preference. It consists of 137k pairs of expert comparisons, based on text prompts and corresponding model outputs from DiffusionDB. To build the ImageRewadDB, we design a pipeline tailored for it, establishing criteria for quantitative assessment and annotator training, optimizing labeling experience, and ensuring quality validation. \mmlu-random-2mmlu-random-1LongWriter-6k
LongWriter-6k
🤗 [LongWriter Dataset] • 💻 [Github Repo] • 📃 [LongWriter Paper]
LongWriter-6k dataset contains 6,000 SFT data with ultra-long output ranging from 2k-32k words in length (both English and Chinese). The data can support training LLMs to extend their maximum output window size to 10,000+ words.
All Models
We open-sourced the following list of models trained on LongWriter-6k:
Model
Huggingface Repo
Description
LongWriter-glm4-9b
🤗… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongWriter-6k.LVBench
LVBench: An Extreme Long Video Understanding Benchmark
[🍎 Project Page] [📖 arXiv Paper] [📊 Dataset][🏆 Leaderboard]
LVBench is a benchmark designed to evaluate and enhance the capabilities of multimodal models in understanding and
extracting information from long videos up to two hours in duration.
🔥 News
2024.06.11 🌟 We released LVBench, a new benchmark for long video understanding!
👀 Introduce to LVBench
LVBench is a benchmark designed to… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LVBench.SWE-Dev-train📝 Paper | 🌐 Github
🤗 SWE-Dev-7B (Qwen-2.5-Coder-7B-Instruct)
🤗 SWE-Dev-9B (GLM-4-9B-Chat)
🤗 SWE-Dev-32B (Qwen-2.5-Coder-32B-Instruct)
🤗 SWE-Dev-train (Training Data)
🚀 SWE-Dev, an open-source Agent for Software Engineering tasks! This repository contains the SWE-Dev-32B model as presented in the paper SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling.
💡 We develop a comprehensive pipeline for creating developer-oriented datasets from GitHub… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/SWE-Dev-train.Vision2Web
Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
[🏠 Project Page] [📖 arXiv Paper] [🏆 Leaderboard] [📮 Submit Results]
Vision2Web is a comprehensive benchmark designed to evaluate multimodal coding agents on visual website development tasks spanning the full software development lifecycle.
This dataset repository contains the benchmark tasks, UI prototypes, test workflows, and resources used to evaluate agent performance.… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/Vision2Web.mmlu-random-Dfineweb2-arb-eduZainAftab-OSworld-Hglm-simple-evals-dataset
glm-simple-evals-dataset
This repository is dedicated to storing various evaluation data required for the glm-simple-evals evaluation project, to enable industry researchers and developers to reproduce the performance of the GLM-4.5 series models on reported benchmarks.
Currently, this repository covers the data required for the following evaluation tasks:
AIME
GPQA
HLE
LiveCodeBench
MATH 500
SciCode
MMLU Pro
Usage Instructions
To use these evaluation datasets… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/glm-simple-evals-dataset.LongCite-45k
LongCite-45k
🤗 [LongCite Dataset] • 💻 [Github Repo] • 📃 [LongCite Paper]
LongCite-45k dataset contains 44,600 long-context QA instances paired with sentence-level citations (both English and Chinese, up to 128,000 words). The data can support training long-context LLMs to generate response and fine-grained citations within a single output.
Data Example
Each instance in LongCite-45k consists of an instruction, a long context (divided into sentences), a user… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongCite-45k.coqa_expanded\\nCoQA: A Conversational Question Answering ChallengeCC-Bench-trajectories
CC-Bench Trajectories Overview
To evaluate GLM-4.6's agentic coding capabilities in real-world scenarios, we developed CC-Bench-V1.1 using Claude Code as the agentic coding testbed. Building on CC-Bench-V1.0, we added 22 more challenging coding tasks and conducted comprehensive evaluations against Claude-Sonnet-4, GLM-4.5, Kimi-K2-0905, and DeepSeek-V3.1-Terminus. The benchmark comprises 74 coding tasks spanning frontend development, tool development, data analysis, testing, and… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/CC-Bench-trajectories.quac_expanded\\nQuestion Answering in Context is a dataset for modeling, understanding,
and participating in information seeking dialog. Data instances consist
of an interactive dialog between two crowd workers: (1) a student who
poses a sequence of freeform questions to learn as much as possible
about a hidden Wikipedia text, and (2) a teacher who answers the questions
by providing short excerpts (spans) from the text. QuAC introduces
challenges not found in existing machine comprehension datasets: its
questions are often more open-ended, unanswerable, or only meaningful
within the dialog context.ComplexFuncBench
Introduction
Complex Function Calling Benchmark (ComplexFuncBench) is specillly designed for complex function calling evaluation. The ComplexFuncBench dataset encompass 1,000 complex function calling samples from five aspects: (1) Function calling with multi-step in single turn; (2) Function calling with user-provided constraints; (3) Function calling that requires parameter value reasoning from implicit information; (4) Function calling with long parameter values that exceed 500… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/ComplexFuncBench.
