datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
terminal-bench-2.0Warning: The leaderboard above is unofficial. The official leaderboard is https://www.tbench.ai/leaderboard/terminal-bench/2.0, in which entires are audited for correct configuration, results show which agent harness is used, and verified trajectories are publicly viewable.
Warning: The dataset is a read-only mirror. The primary source for this dataset is on GitHub: https://github.com/harbor-framework/terminal-bench-2. Please open issues and pull requests there.
How this mirror was created… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-2.0.big_bench_hardAll rights and obligations of the dataset are with original authors of the paper/dataset.
I have merely made this dataset with a MIT licence available on HuggingFace.
BIG-Bench Hard Dataset
This repository contains a copy of the BIG-Bench Hard dataset.
Small edits to the formatting of the dataset are made to integrate it into the Inspect Evals repository, a community contributed LLM
evaulations for Inspect AI a framework by the UK AI Safety Institute.
The BIG-Bench Hard dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/Joschka/big_bench_hard.AmericanStoriesAmerican Stories offers high-quality structured data from historical newspapers suitable for pre-training large language models to enhance the understanding of historical English and world knowledge. It can also be integrated into external databases of retrieval-augmented language models, enabling broader access to historical information, including interpretations of political events and intricate details about people's ancestors. Additionally, the structured article texts facilitate the application of transformer-based methods for popular tasks like detecting reproduced content, significantly improving accuracy compared to traditional OCR methods. American Stories serves as a substantial and valuable dataset for advancing multimodal layout analysis models and other multimodal applications.BeyondSWE-harbor
BeyondSWE-harbor
This repository provides the harbor version of the BeyondSWE benchmark, containing the full task instances in a directory-based (harbor) format, where each instance is stored as an independent folder.
📌 For benchmark definition, data format and detailed evaluation results, please refer to: 🤗 Main Dataset on HuggingFace
🗂️ Data Structure
beyondswe/
├── {instance_id}/
│ ├── environment/
│ ├── solution/
│ ├── tests/
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/AweAI-Team/BeyondSWE-harbor.kernelbench-hard-traces
KernelBench-Hard agent traces
Frontier coding agents writing optimized CUDA/Triton kernels (FP8 GEMM, paged
attention, MoE, W4A16, KDA, Top-k) on RTX PRO 6000 Blackwell, H100 PCIe, and
B200; roofline-graded.
Each .jsonl file is one agent run in Claude-Code session format, viewable with
the Hugging Face Agent Trace viewer (Data Studio → open a row). Filename =
run id.
Live leaderboard: https://kernelbench.com/hard
Secrets redacted. Full reasoning for open-provider routes… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-traces.harness-1-train-data
Harness-1 Training Data
This dataset contains the training data used for Harness-1, plus the retrieval
corpora needed to reproduce the training/evaluation environment.
Contents
The dataset has one train split with a stage column:
sft: 899 raw GPT-5.4-generated v8d SFT trajectories produced by generate_sft_ultra_0417.py and used by train_sft_ultra_0417.py from sft_ultra_v8d_data.
rl: 3453 SEC training-split query records used for RL (TRAIN_DATASETS=sec… See the full description on the dataset page: https://huggingface.co/datasets/pat-jj/harness-1-train-data.newswire
Dataset Card for NewsWire
Dataset Summary
NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model.
Languages
English (en)
Dataset Structure
Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.scorio-lite
Scorio Lite contains 1,211,520 sampled attempts from four model configurations and
six reasoning benchmarks. Each model was run 80 times on every question.
The five competition-math splits contain 186 questions. The superGPQA split contains a
frozen, field-balanced sample of 3,600 questions. Each row includes the generation,
rule-based grading, scores from CompassVerifier-3B and a reference-free verifier, and
aggregate token statistics. Per-model configs also include token strings, log… See the full description on the dataset page: https://huggingface.co/datasets/harimo/scorio-lite.scorio-trace
Scorio Trace contains 192,000 sampled reasoning traces from 20 model configurations and
four competition math benchmarks. Each model was run 80 times on each of the 30 questions
in every benchmark.
Each row contains one complete generation, its rule-based correctness, scores from two
reward models, and token-level log probabilities and vocabulary ranks. The 80 generations
for one model, task, and question form a candidate pool. They are ordered by seed, so
pool[:n] gives a reproducible sample… See the full description on the dataset page: https://huggingface.co/datasets/harimo/scorio-trace.chrono-2020-harvest
ChronoLLM 2020 harvest
Public gated (gated=manual) dataset of year-cutoff training text for Bittensor subnet 38. Calendar window 2020-01-01 … 2020-12-31.
Repo: jjjlimaus/chrono-2020-harvest
Last sync (UTC): 2026-09-04T18:45:47.313133+00:00
Request access on Hugging Face and wait for manual review. Do not treat this as a license to republish every source; some feeds are public-domain, others are news/API text kept for research training only.
Streaming layout… See the full description on the dataset page: https://huggingface.co/datasets/jjjlimaus/chrono-2020-harvest.harbor-mix
Harbor-Mix
Harbor-Mix is a curated meta-dataset of 100 difficult, diverse, and high-quality agentic evaluation tasks selected from the Harbor Adapters benchmark pool. It is designed to preserve broad signal from large-scale agent evaluations while being substantially cheaper to run than a full multi-benchmark sweep.
What Is Included
This repository contains the 100 task directories, flattened at the repository root.
Each task directory contains:
instruction.md:… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/harbor-mix.HarmfulQAPaper | Github | Dataset| Model
📣📣📣: Do check our new multilingual dataset CatQA here used in Safety Vectors:📣📣📣
As a part of our research efforts toward making LLMs more safe for public use, we create HarmfulQA i.e. a ChatGPT-distilled dataset constructed using the Chain of Utterances (CoU) prompt. More details are in our paper Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment
HarmfulQA serves as both-a new LLM safety benchmark and an alignment dataset… See the full description on the dataset page: https://huggingface.co/datasets/declare-lab/HarmfulQA.tts-datagen
GPT-OSS 120B native reasoning traces for TTS Datagen
Summary
This dataset contains 2,865 synthetic competitive-programming questions,
45,840 independently sampled GPT-OSS 120B solutions (16 per question), and 50
verified test cases per question (143,250 test cases total). Each solution
preserves the model's native reasoning trace separately from its final answer.
The reasoning was returned by MetaGen's native Dialog Completion interface as
dialog reasoning… See the full description on the dataset page: https://huggingface.co/datasets/harman/tts-datagen.terminal-bench-2-verified
Terminal-Bench 2.0 Verified: Instruction & Environment Fix Version
中文版本
We conducted a comprehensive review of the entire Terminal-Bench 2.0 dataset and identified various issues. Both GLM-5 and Step 3.5-Flash were evaluated using this verified version.
This modified version addresses environment and instruction issues we discovered in Terminal-Bench 2.0. It includes two types of fixes:
Environment Fixes: Updated Dockerfiles and instructions to support Claude Code Agent runtime… See the full description on the dataset page: https://huggingface.co/datasets/harithoppil/terminal-bench-2-verified.dart-math-hard
🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub
🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX
[!IMPORTANT]
🔥 Excited to find our DART-Math-DSMath-7B (Prop2Diff) trained on DART-Math-Hard comparable to the AIMO winner NuminaMath-7B on CoT,
but based solely on MATH & GSM8K prompt set, leaving much room to improve!
Besides, our DART method is also fully compatible… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/dart-math-hard.chrono-2021-harvest
ChronoLLM 2021 harvest
Public gated (gated=manual) dataset of year-cutoff training text for Bittensor subnet 38. Calendar window 2021-01-01 … 2021-12-31.
Repo: jjjlimaus/chrono-2021-harvest
Last sync (UTC): 2026-09-12T17:44:33.508539+00:00
Request access on Hugging Face and wait for manual review. Do not treat this as a license to republish every source; some feeds are public-domain, others are news/API text kept for research training only.
Streaming layout… See the full description on the dataset page: https://huggingface.co/datasets/jjjlimaus/chrono-2021-harvest.hardtests_problems
Dataset Card for HARDTESTS Problems
HARDTESTS is a competitive programming dataset containing 47,136 problems collected from 13 different Online Judges (OJs). Each problem includes a problem statement, numerous oracle code solutions, and a set of relatively reliable test cases. Note: Due to their large size, the test cases are stored in a separate dataset. This dataset is presented in the paper HardTests: Synthesizing High-Quality Test Cases for LLM Coding.
Project Page… See the full description on the dataset page: https://huggingface.co/datasets/sigcp/hardtests_problems.harbor-swesmith-rl-artifacts
Harbor SWE-Smith 强化学习数据产物
本数据集是 Harbor Qwen 工具调用代码智能体强化学习项目使用的冻结任务集,服务于 GRPO、原生价值模型/GAE PPO、训练过程诊断和统一协议评测。
项目已于 2026 年 8 月 30 日完成 P0 评测并进入阶段性归档。本数据集用于保留实验所依赖的数据切分、任务执行文件和审计信息,不代表新的通用代码能力基准。
数据概况
切分
任务数
训练集
187
验证集
42
测试集
38
合计
267
数据覆盖 89 个上游代码仓库。三个切分之间同时执行任务标识和仓库级隔离检查。
正式数据集名称:
swesmith-curated-grpo-267-v1
冻结切分的语义摘要:
ae5df9a3f4a3fc8af44fac420b36529e283839e1bd3de9daba65d5bcda51447d
该值来自 split-manifest.json 的 sha256 字段,用于标识切分语义,不等同于该文件本身的字节级… See the full description on the dataset page: https://huggingface.co/datasets/keryszhan/harbor-swesmith-rl-artifacts.glm53-flash-harvest
GLM-5.3-Flash On-Policy Harvest
86,006 responses / 246,034,910 generated tokens written by
zai-org/GLM-5.3-Flash from its reference FP8 weights,
across four harvest rounds, 15 registers and both serving modes (22,016 rows carry the
model's inline <think>…</think> chain). It is on-policy text: the corpus records what the target model
actually generates, which is what a speculative-decoding drafter (EAGLE-3 / DFlash / DSpark family) has to
learn to predict. Everything here is MIT.… See the full description on the dataset page: https://huggingface.co/datasets/Zek-Takai/glm53-flash-harvest.tts
TTS synthetic programming-question dataset
This public dataset contains the frozen set of 2,865 accepted programming
questions and their materialized verifier tests.
Accepted question bundle
accepted_bundle/accepted-questions-2865.tar.zst contains all 42,975 accepted
question files: statements, package JSON, three public examples, generators,
validators, reference solutions, brute-force solutions, verification records,
provenance records, and GPT-OSS hardness… See the full description on the dataset page: https://huggingface.co/datasets/harman/tts.hardtests_tests
Dataset Card for HARDTESTS Tests
HARDTESTS Tests is the test suite of HARDTESTS, a competitive programming dataset. The dataset contains multiple .parquet files. Each instance contains structured test case objects for one problem. This dataset is presented in the paper HardTests: Synthesizing High-Quality Test Cases for LLM Coding.
Project Page
Data Summary
The test suite is generated using the HARDTESTSGEN pipeline.
The dataset contains the generated test suites of… See the full description on the dataset page: https://huggingface.co/datasets/sigcp/hardtests_tests.LifelongSAThis is the two iteration defender of NeurIPS 2025 "Lifelong Safety Alignment for Language Models": https://openreview.net/forum?id=9YkEcAqiIK
The defenders are trained on RR and LAT.
LEXam-hard
LEXam-hard
The 518 open questions of LEXam
that the strongest open models score lowest on. LEXam is a benchmark of law exam
questions from the University of Zurich (Fan et al., ICLR 2026;
website, code).
This is a filtered copy of its open_question test split for evaluating agents that
can do legal research. Questions that strong models already answer from memory are
removed, and so are questions that cannot be answered from their own text.
Rows are ordered from the hardest… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/LEXam-hard.DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
DeepSeek v4 Pro Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-pro.
JSONL files: 4006
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/hardcoremoore/DeepSeek-v4-Pro-Agent.chrono-2022-harvest
ChronoLLM 2022 harvest
Public gated (gated=manual) dataset of year-cutoff training text for Bittensor subnet 38. Calendar window 2022-01-01 … 2022-12-31.
Repo: jjjlimaus/chrono-2022-harvest
Last sync (UTC): 2026-09-17T12:52:33.701253+00:00
Request access on Hugging Face and wait for manual review. Do not treat this as a license to republish every source; some feeds are public-domain, others are news/API text kept for research training only.
Streaming layout… See the full description on the dataset page: https://huggingface.co/datasets/jjjlimaus/chrono-2022-harvest.chrono-2021-quality-harvest
ChronoLLM 2021 quality cutoff corpus
Clean, cross-source-deduplicated training text available no later than
2021-12-31. Target: 15B exact anacoluthe89/chrono-2015 tokens. The build rejects
any row that would exceed 15B; 20B is an external safety ceiling, not a
collection goal.
Repository: jjjlimaus/chrono-2021-quality-harvest
Initialized: 2026-09-17T13:43:12.278373+00:00
Mix and hard quotas
Lane
Maximum tokens
Cutoff basis
English Wikipedia snapshot… See the full description on the dataset page: https://huggingface.co/datasets/jjjlimaus/chrono-2021-quality-harvest.harvey-notes-v4
wm-rl notes v4 — two note banks from a recursive self-experience loop
Continuation update (2026-09-14): rounds 6–15 appended. The recursive bank now contains 308,580 notes / 138,723,516 training tokens. The original 108,081-row round-5 bank remains an exact prefix. Round-4/5 task lists were available for duplicate rejection, but their trajectory manifests were not published; the continuation therefore seeds round 6 from the published recall sessions and carries complete… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-notes-v4.volume2gym-railroad-1959
The Rulebook Becomes a World
Volume2Gym asks a deliberately expansive question: what if any sufficiently structured text could become a small, inspectable world in which a model learns by acting, receiving feedback, and trying again?
This release turns one bounded English technical volume into an auditable reinforcement-learning dataset: 117 source scans → 536 extracted rules → 2,708 synthetic scenario tasks → a measured rule-linkage and verification surface. It is an… See the full description on the dataset page: https://huggingface.co/datasets/HarleyCooper/volume2gym-railroad-1959.gemini-3.1-pro-hard-high-reasoning
Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M
Dataset Details
Dataset Description
This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification.
The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3.1-pro-hard-high-reasoning.DeepScaleR-Easy-Medium-Hard-Gemma-26B-PT-10k
DeepScaleR Easy/Medium/Hard — Gemma 4 26B-A4B PT
This dataset contains 9,900 unique, deduplicated DeepScaleR math questions for
reinforcement-learning experiments. Difficulty is defined by how often the
pretrained google/gemma-4-26B-A4B teacher solved each question across eight
temperature-1 samples under the same rule-based grader used by the RL training
pipeline.
The Hub dataset has three configurations—easy, medium, and hard—and each
configuration has a train split with 3,000… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DeepScaleR-Easy-Medium-Hard-Gemma-26B-PT-10k.
