datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kernelbench-hard-traces
KernelBench-Hard agent traces
Frontier coding agents writing optimized CUDA/Triton kernels (FP8 GEMM, paged
attention, MoE, W4A16, KDA, Top-k) on RTX PRO 6000 Blackwell, H100 PCIe, and
B200; roofline-graded.
Each .jsonl file is one agent run in Claude-Code session format, viewable with
the Hugging Face Agent Trace viewer (Data Studio → open a row). Filename =
run id.
Live leaderboard: https://kernelbench.com/hard
Secrets redacted. Full reasoning for open-provider routes… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-traces.newswire
Dataset Card for NewsWire
Dataset Summary
NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model.
Languages
English (en)
Dataset Structure
Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.scorio-lite
Scorio Lite contains 1,211,520 sampled attempts from four model configurations and
six reasoning benchmarks. Each model was run 80 times on every question.
The five competition-math splits contain 186 questions. The superGPQA split contains a
frozen, field-balanced sample of 3,600 questions. Each row includes the generation,
rule-based grading, scores from CompassVerifier-3B and a reference-free verifier, and
aggregate token statistics. Per-model configs also include token strings, log… See the full description on the dataset page: https://huggingface.co/datasets/harimo/scorio-lite.scorio-trace
Scorio Trace contains 192,000 sampled reasoning traces from 20 model configurations and
four competition math benchmarks. Each model was run 80 times on each of the 30 questions
in every benchmark.
Each row contains one complete generation, its rule-based correctness, scores from two
reward models, and token-level log probabilities and vocabulary ranks. The 80 generations
for one model, task, and question form a candidate pool. They are ordered by seed, so
pool[:n] gives a reproducible sample… See the full description on the dataset page: https://huggingface.co/datasets/harimo/scorio-trace.tts-datagen
GPT-OSS 120B native reasoning traces for TTS Datagen
Summary
This dataset contains 2,865 synthetic competitive-programming questions,
45,840 independently sampled GPT-OSS 120B solutions (16 per question), and 50
verified test cases per question (143,250 test cases total). Each solution
preserves the model's native reasoning trace separately from its final answer.
The reasoning was returned by MetaGen's native Dialog Completion interface as
dialog reasoning… See the full description on the dataset page: https://huggingface.co/datasets/harman/tts-datagen.harbor-swesmith-rl-artifacts
Harbor SWE-Smith 强化学习数据产物
本数据集是 Harbor Qwen 工具调用代码智能体强化学习项目使用的冻结任务集,服务于 GRPO、原生价值模型/GAE PPO、训练过程诊断和统一协议评测。
项目已于 2026 年 8 月 30 日完成 P0 评测并进入阶段性归档。本数据集用于保留实验所依赖的数据切分、任务执行文件和审计信息,不代表新的通用代码能力基准。
数据概况
切分
任务数
训练集
187
验证集
42
测试集
38
合计
267
数据覆盖 89 个上游代码仓库。三个切分之间同时执行任务标识和仓库级隔离检查。
正式数据集名称:
swesmith-curated-grpo-267-v1
冻结切分的语义摘要:
ae5df9a3f4a3fc8af44fac420b36529e283839e1bd3de9daba65d5bcda51447d
该值来自 split-manifest.json 的 sha256 字段,用于标识切分语义,不等同于该文件本身的字节级… See the full description on the dataset page: https://huggingface.co/datasets/keryszhan/harbor-swesmith-rl-artifacts.glm53-flash-harvest
GLM-5.3-Flash On-Policy Harvest
86,006 responses / 246,034,910 generated tokens written by
zai-org/GLM-5.3-Flash from its reference FP8 weights,
across four harvest rounds, 15 registers and both serving modes (22,016 rows carry the
model's inline <think>…</think> chain). It is on-policy text: the corpus records what the target model
actually generates, which is what a speculative-decoding drafter (EAGLE-3 / DFlash / DSpark family) has to
learn to predict. Everything here is MIT.… See the full description on the dataset page: https://huggingface.co/datasets/Zek-Takai/glm53-flash-harvest.DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
DeepSeek v4 Pro Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-pro.
JSONL files: 4006
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/hardcoremoore/DeepSeek-v4-Pro-Agent.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 5.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 1.3000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 4.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think.swerl-tmax-15k-hardened-prefilter
swerl-tmax-15k hardened, pre-validation-filter (dataset 2 of 3)
The middle artifact of three, which exist to be compared against each other by
task_id:
original — hamishivi/swerl-tmax-15k, unchanged.
hardened, pre-validation-filter — this dataset.
hardened, post-validation-filter — wAI-org/swerl-tmax-15k-solvable-gpt-5-6-terra, a strict subset of this one — 7,015 tasks.
What is in it
Tasks where a patch was actually applied, plus tasks originally labelled
CLEAN… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/swerl-tmax-15k-hardened-prefilter.seas-harden-only
seas-harden-only
Per-round trajectory + per-category taxonomy from an automated red-teaming co-evolution loop (attacker -> target -> harm-judge) on Qwen2.5 (7B attacker+judge, 3B target), seeded with JailbreakBench behaviors. This dataset is the frozen-attacker control, canned refusal (only the target hardens) arm: held-in ASR flat ~0.45 across 5 rounds - LoRA-DPO hardening toward a canned refusal does not reduce ASR.
Study question: does adversarial co-evolution ignite an arms… See the full description on the dataset page: https://huggingface.co/datasets/yavuz-ai/seas-harden-only.harvey-lab-glm-traces
Harvey LAB teacher traces (GLM-5.2 → Qwen3.5-9B SFT)
Agentic tool-use traces collected from a GLM-5.2 teacher solving Harvey LAB
legal benchmark tasks inside the open-rl scaffold (bash / read / write / todo tools,
sandboxed workspace, 163,840-token trajectory budget, 32k max tokens per turn).
v2: the teacher's chain-of-thought is captured per turn in the reasoning
field (--reasoning-parser on the serving endpoint), so students can be trained
to think before acting — SFT on the… See the full description on the dataset page: https://huggingface.co/datasets/ShubyM/harvey-lab-glm-traces.taskweft-fbd-harness-train
taskweft-fbd-harness-train
Intents and the IEC 61131-3 Function Block Diagrams that carry them out, as an
EditScore-shaped corpus: one root row per intent, three candidates per row (rank1 the
reference diagram, rank3 one that compiles and does the wrong thing, rank5 one the
compiler refuses), and one score row per candidate from the compiler's reference scan, rank1's outputs the reference for the others. Every row is
constructed from a template and a seed, so the labels are true… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/taskweft-fbd-harness-train.Olympiads_hard
Numina-Olympiads
Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers.
Dataset Information
Split: train
Original size: 21525
Filtered size: 21408
Source: olympiads
All examples contain valid boxed answers
Dataset Description
This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes:
A mathematical word problem
A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Olympiads_hard.nb-asr-numerics-harvested
Norwegian Bokmål Numeric Expression Harvesting Dataset
This dataset contains cleaned, high-quality Norwegian Bokmål sentences containing numeric expressions harvested from both the Norwegian Colossal Corpus (NbAiLab/NCC) and the FineWeb-2 Norwegian Bokmål subset (HuggingFaceFW/fineweb-2 config nob_Latn).
This is a combined high-volume intermediate dataset built for the first stage of a template-based Norwegian synthetic speech (TTS) generation pipeline to improve number/digit… See the full description on the dataset page: https://huggingface.co/datasets/pere/nb-asr-numerics-harvested.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 2.2000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think.harmonic-reasoning-v1
Harmonic Reasoning v1
Support This Work
I'm a PhD student in visual neuroscience at the University of Toronto who also happens to spend way too much time fine-tuning, merging, and quantizing open-weight models on rented H100s and a local DGX Spark. All training compute is self-funded — balancing GPU costs against a student budget. If my open-weight models or datasets have been useful to you, consider supporting future releases.
Support on Ko-fi
Harmonic Reasoning v1 is a… See the full description on the dataset page: https://huggingface.co/datasets/DJLougen/harmonic-reasoning-v1.hsk30-graded-readers
HSK 3.0 Graded Reader Corpus
132 word-aligned Chinese graded readers with per-word pinyin and English gloss,
arranged on six difficulty shelves: 102 texts in the main split and a
disjoint 30-text held-out split.
Aligned Chinese graded-reader corpora are scarce. Existing collections are
unaligned plain text, locked inside commercial apps, or graded against HSK 2.0,
which has been superseded twice.
Loading
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/harukicoder/hsk30-graded-readers.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.05.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.01.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.1.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-10m-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-10m-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 10M notes + note-conditioned trajectory mixture, and no KL regularization.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-10m-think.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-1m-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-1m-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 1M notes + note-conditioned trajectory mixture, and no KL regularization.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-1m-think.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-30m-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-30m-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 30M notes + note-conditioned trajectory mixture, and no KL regularization.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-30m-think.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and no KL regularization.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think.data_agent_harbor_train_sft
data_agent_harbor_train_sft
4677 verified agent trajectories (SFT) for the data_agent_harbor_train environments. TRL-ready tool-calling format: messages + tools columns.
Each row is a reward=1 rollout — instruction -> bash tool calls (shell commands) -> final answer — graded deterministically (no LLM judge). Single bash tool throughout.
Columns
messages: OpenAI/TRL chat format (system, user, assistant+tool_calls, tool, ...). tool_calls[].function.arguments are… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/data_agent_harbor_train_sft.harvey-eval-gpt56sol-qwen35-9b-base-20t-think
harvey-eval-gpt56sol-qwen35-9b-base-20t-think
4,000 evaluation attempts: 250 tasks × 16 samples. Original samples 0–3 and
twelve additional seeded samples 4–15. The train split contains evaluation
records, not training data. The evaluated model is Qwen/Qwen3.5-9B at
revision c202236235762e1c871ad0ccb60c8ee5ba337b9a.
Cohort
Attempts
All-criteria-pass rate ± task-level SEM
Original four
1,000
3.900% ± 0.986 pp
Additional twelve
3,000
3.533% ± 0.882 pp
Combined… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-base-20t-think.Polaris-hard-w-solutions-24209
Polaris-Hard-w-Solutions
24,209 hard competition-math problems (the hardest difficulty bands of the
Polaris dataset) paired with
two verified solutions each: a full original solution and a concise summarized solution. Every
retained problem has a machine-verifiable final answer, every solution's boxed answer grades correct
against the reference (sympy-based grading), and the summarized solutions have additionally been put
through a reasoning-rigor pass (see step 5 below).
This… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/Polaris-hard-w-solutions-24209.
