CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zai-org /LongBenchLongBench is a comprehensive benchmark for multilingual and multi-task purposes, with the goal to fully measure and evaluate the ability of pre-trained language models to understand long text. This dataset consists of twenty different tasks, covering key long-text application scenarios such as multi-document QA, single-document QA, summarization, few-shot learning, synthetic tasks, and code completion.question-answering1K<n<10K191 likes59k downloads2y agoHugging Face02zai-org /humaneval-xHumanEval-X is a benchmark for the evaluation of the multilingual ability of code generative models. It consists of 820 high-quality human-crafted data samples (each with test cases) in Python, C++, Java, JavaScript, and Go, and can be used for various tasks.text-generation97 likes2.5k downloads4y agoHugging Face03zai-org /LongWriter-6k LongWriter-6k 🤗 [LongWriter Dataset] • 💻 [Github Repo] • 📃 [LongWriter Paper] LongWriter-6k dataset contains 6,000 SFT data with ultra-long output ranging from 2k-32k words in length (both English and Chinese). The data can support training LLMs to extend their maximum output window size to 10,000+ words. All Models We open-sourced the following list of models trained on LongWriter-6k: Model Huggingface Repo Description LongWriter-glm4-9b 🤗… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongWriter-6k.texttext-generation1K<n<10K205 likes1.7k downloads2y agoHugging Face04zai-org /Vision2Web Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification [🏠 Project Page] [📖 arXiv Paper] [🏆 Leaderboard] [📮 Submit Results] Vision2Web is a comprehensive benchmark designed to evaluate multimodal coding agents on visual website development tasks spanning the full software development lifecycle. This dataset repository contains the benchmark tasks, UI prototypes, test workflows, and resources used to evaluate agent performance.… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/Vision2Web.imagetext-generationn<1K23 likes1.4k downloads6mo agoHugging Face05zai-org /LongCite-45k LongCite-45k 🤗 [LongCite Dataset] • 💻 [Github Repo] • 📃 [LongCite Paper] LongCite-45k dataset contains 44,600 long-context QA instances paired with sentence-level citations (both English and Chinese, up to 128,000 words). The data can support training long-context LLMs to generate response and fine-grained citations within a single output. Data Example Each instance in LongCite-45k consists of an instruction, a long context (divided into sentences), a user… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongCite-45k.texttext-generation10K<n<100K78 likes614 downloads2y agoHugging Face06zai-org /CC-Bench-trajectories CC-Bench Trajectories Overview To evaluate GLM-4.6's agentic coding capabilities in real-world scenarios, we developed CC-Bench-V1.1 using Claude Code as the agentic coding testbed. Building on CC-Bench-V1.0, we added 22 more challenging coding tasks and conducted comprehensive evaluations against Claude-Sonnet-4, GLM-4.5, Kimi-K2-0905, and DeepSeek-V3.1-Terminus. The benchmark comprises 74 coding tasks spanning frontend development, tool development, data analysis, testing, and… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/CC-Bench-trajectories.tabulartext-generationn<1K98 likes510 downloads1y agoHugging Face07zai-org /LongReward-10k LongReward-10k 💻 [Github Repo] • 📃 [LongReward Paper] LongReward-10k dataset contains 10,000 long-context QA instances (both English and Chinese, up to 64,000 words). The sft split contains SFT data generated by GLM-4-0520, following the self-instruct method in LongAlign. Using this split, we supervised fine-tune two models: LongReward-glm4-9b-SFT and LongReward-llama3.1-8b-SFT, which are based on GLM-4-9B and Meta-Llama-3.1-8B, respectively. The dpo_glm4_9b and… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongReward-10k.texttext-generation10K<n<100K8 likes375 downloads2y agoHugging Face08zait-ai /RuHeritage-Corpus RuHeritage-Corpus 🇬🇧 English Description RuHeritage-Corpus is a high-quality, curated dataset of Russian classical literature, specifically designed for the pre-training and continued pre-training (CPT) of Large Language Models (LLMs). The corpus focuses on the Golden and Silver Ages of Russian literature, providing models with exposure to rich vocabulary, complex syntactic structures, and stylistically flawless Russian text, acting as a "quality anchor"… See the full description on the dataset page: https://huggingface.co/datasets/zait-ai/RuHeritage-Corpus.tabulartext-generation1K<n<10K2 likes362 downloads2mo agoHugging Face09zai-org /webglm-qa WebGLM-QA Dataset Description WebGLM-QA is the dataset used to train the WebGLM generator module. It consists of 43,579 high-quality data samples for the train split, 1,000 for the validation split, and 400 for the test split. Refer to our paper for the data construction details. Dataset Structure To load the dataset, you can try the following code. from datasets import load_dataset load_dataset("THUDM/webglm-qa") DatasetDict({ train: Dataset({… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/webglm-qa.texttext-generation10K<n<100K65 likes361 downloads3y agoHugging Face10zai-org /ZClawBench ZClawBench Overview Recent advances in agent frameworks have pushed large language models beyond conversational assistance toward goal-driven task execution. Among these frameworks, OpenClaw has emerged as a representative setting for evaluating whether models can interact with tools, follow multi-step instructions, and complete practical tasks in realistic environments. Unlike traditional chatbot benchmarks, OpenClaw-style scenarios require models not only to produce… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/ZClawBench.texttext-generationn<1K34 likes221 downloads6mo agoHugging Face11zai-org /BPO Dataset Card for Black-box Prompt Optimization (BPO) Data Summary To advance the development of alignment in language models, we introduce a black-box alignment method. BPO enhances the alignment of various Large Language Models (LLMs) with human preferences using only a plug-and-play model. To further promote alignment work from the prompting perspective, we are releasing the BPO Dataset. This dataset comprises 14,395 entries of prompt optimization pairs, constructed… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/BPO.texttext-generation10K<n<100K24 likes141 downloads3y agoHugging Face12zait-ai /OpenJA OpenJA 🇬🇧 English Description OpenJA is a dataset of clean, officially published parliamentary transcripts of the Japanese language. It is characterized by high-quality text (without web noise, HTML, advertising) and reliable metadata, but it represents one narrow language register (official/parliamentary speech) and is better suited as an addition to more diverse corpora than as the only source for a general-purpose pretrain. 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/zait-ai/OpenJA.tabulartext-generation1M<n<10M1 likes130 downloads2mo agoHugging Face13zai-org /SurveyReview SurveyReview SurveyReview is a reviewer-aligned benchmark for evaluating survey papers. It turns real peer-review reports into multidimensional scores and rationales so that model judgments can be compared with human reviewer judgments. The benchmark covers four dimensions: Readability, Criticalness, Comprehensiveness, and Structure. Latest release: v1.1 What's New in v1.1 Cleaned full-text content for 1,646 survey articles, stored in two JSON shards. The… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/SurveyReview.imagetext-classification0 likes118 downloads2mo agoHugging Face14zait-ai /OpenOcypus-1.0 🪲 OpenOcypus-1.0 📖 Description OpenOcypus-1.0 is the first release in the OpenOcypus series — a collection of high‑quality SFT datasets designed to evolve over time. Future versions will introduce additional sources, refined filtering, and expanded task coverage. Total size: 1,157,428 examples. The dataset is designed to create a versatile assistant capable of: 🗣️ Engaging in natural conversations 🧮 Solving math problems with step‑by‑step explanations 💻… See the full description on the dataset page: https://huggingface.co/datasets/zait-ai/OpenOcypus-1.0.texttext-generation1M<n<10M0 likes108 downloads23d agoHugging Face15zaibutcooler /mini-burmesetexttext-generation100K<n<1M1 likes39 downloads1y agoHugging Face16zai-org /Vision2Web-Leaderboard Vision2Web Leaderboard Submissions This repository accepts leaderboard submissions for Vision2Web, a benchmark for visual website development agents. Submissions should contain the inference outputs (i.e., the generated code/files) for each task. Evaluation is conducted by the maintainers using the latest VLM Judge and GUI Agent, ensuring fair and consistent scoring across all submissions. Seasons Vision2Web Leaderboard is organized into seasons, each lasting 3… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/Vision2Web-Leaderboard.text-generation4 likes29 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.