datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kernelbench-hard-traces
KernelBench-Hard agent traces
Frontier coding agents writing optimized CUDA/Triton kernels (FP8 GEMM, paged
attention, MoE, W4A16, KDA, Top-k) on RTX PRO 6000 Blackwell, H100 PCIe, and
B200; roofline-graded.
Each .jsonl file is one agent run in Claude-Code session format, viewable with
the Hugging Face Agent Trace viewer (Data Studio → open a row). Filename =
run id.
Live leaderboard: https://kernelbench.com/hard
Secrets redacted. Full reasoning for open-provider routes… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-traces.newswire
Dataset Card for NewsWire
Dataset Summary
NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model.
Languages
English (en)
Dataset Structure
Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.imagenet_hard_review_data_r2harbor-swesmith-rl-artifacts
Harbor SWE-Smith 强化学习数据产物
本数据集是 Harbor Qwen 工具调用代码智能体强化学习项目使用的冻结任务集,服务于 GRPO、原生价值模型/GAE PPO、训练过程诊断和统一协议评测。
项目已于 2026 年 8 月 30 日完成 P0 评测并进入阶段性归档。本数据集用于保留实验所依赖的数据切分、任务执行文件和审计信息,不代表新的通用代码能力基准。
数据概况
切分
任务数
训练集
187
验证集
42
测试集
38
合计
267
数据覆盖 89 个上游代码仓库。三个切分之间同时执行任务标识和仓库级隔离检查。
正式数据集名称:
swesmith-curated-grpo-267-v1
冻结切分的语义摘要:
ae5df9a3f4a3fc8af44fac420b36529e283839e1bd3de9daba65d5bcda51447d
该值来自 split-manifest.json 的 sha256 字段,用于标识切分语义,不等同于该文件本身的字节级… See the full description on the dataset page: https://huggingface.co/datasets/keryszhan/harbor-swesmith-rl-artifacts.kernelbench-hard-runs
KernelBench-Hard — Agent Runs
84 full agent transcripts (12 frontier models × 7 problems) from the KernelBench-Hard sweep on a single Blackwell GPU (RTX PRO 6000, sm_120, CUDA 13.2). Each run contains the model's full reasoning trace, every tool call, the final solution.py, and the eval result.
Companion datasets:
Infatoshi/kernelbench-hard-problems — the 7 problem definitions
Live site: https://kernelbench.com/hard
100 themed transcript viewers (HTML): https://kernelbench.com/runs… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-runs.DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
DeepSeek v4 Pro Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-pro.
JSONL files: 4006
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/hardcoremoore/DeepSeek-v4-Pro-Agent.registry-harvest-xrpl-mica-lei
Free-registry harvest — XRPL issuers × MiCA × LEI
Public registers and public ledgers joined, read 2026-09-04T07:12:30Z. Every source is free, keyless and re-runnable by
anyone. No part of this needed a relationship, an API key, or anyone's permission.
The finding
Of the 16 XRPL issued assets in the CSOAI reader, 5 join to a MiCA e-money-token authorisation.
group
n
declares an on-chain domain
enforces allowlisting
retains freeze capability… See the full description on the dataset page: https://huggingface.co/datasets/csoai/registry-harvest-xrpl-mica-lei.hh-harmless-base-qwen3-8b-margin-dpo-margin-logswikiMIA-2024-hard
WikiMIA-2024 Hard Dataset
Dataset Description
WikiMIA_2024 Hard is a challenging dataset for membership inference attacks intorduced in the paper "The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage" containing temporal Wikipedia articles with different versions based on date cutoffs.
This dataset is designed to evaluate the robustness of privacy-preserving machine learning models against sophisticated membership inference techniques.
It… See the full description on the dataset page: https://huggingface.co/datasets/hallisky/wikiMIA-2024-hard.harbor-release-catalog
February 2025 Harbor Software Catalog
Approved stable releases published in February 2025.
Artifact
Version
Published
Downloads
Buoy Mapper
1.0.0
2025-02-26
5,175
Dock Ledger
3.2.1
2025-02-24
11,980
Harbor Status API
1.4.0
2025-02-14
18,420
Total releases: 3
Total downloads: 35,575
wire_harness_expert_sac
Wire Harness Expert SAC
Expert-policy trajectories collected from the five-mover WireHarness MuJoCo
environment for visual world-model training.
Dataset summary
20,000 episodes
3,491,570 stored observation rows
At most 300 environment transitions per episode (up to 301 stored rows,
including the initial observation)
224 x 224 RGB observations, stored as JPEG bytes in pixels
10-dimensional continuous actions
451-dimensional observations
Five task stages and… See the full description on the dataset page: https://huggingface.co/datasets/faridganbarli/wire_harness_expert_sac.swanlake_hard
swanlake_hard
Synthetic Sokoban hard dataset generated from the local VisGym Sokoban environment.
Contents
trajectories/sokoban_hard/test/*.jsonl
trajectories/sokoban_hard/train/*.jsonl
manifests/
metadata/
Generation Summary
Task: sokoban_hard
Raw test generated: 1200
Final test after dedupe: 1200
Raw train generated: 120000
Final train after dedupe: 112000
Train samples removed by test-hash filter: 0
Train samples removed by train self-dedupe: 8000
Raw… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/swanlake_hard.pacman_hard_cot_chunk_k10_train
pacman_hard_cot_chunk_k10_train
BAGEL VLM-Gym world-model dataset (pacman / cot).
CoT chunk-K train set: all-step interleaved imagined reasoning; re-grounds on the true frame every K=10 steps.
layout: Train-only. Gzipped-JSONL shards under training/; each row is one packed SFT sample with base64-JPEG frames inline.
images are base64-encoded JPEG frames stored inline in each JSONL row.
Pairs with the matching pacman checkpoint(s) under the companion model org; CoT and non-CoT… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/pacman_hard_cot_chunk_k10_train.lm-eval-results-DreadPoor-Harpy-7B-Model_Stock-private
Dataset Card for Evaluation run of DreadPoor/Harpy-7B-Model_Stock
Dataset automatically created during the evaluation run of model DreadPoor/Harpy-7B-Model_Stock
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-DreadPoor-Harpy-7B-Model_Stock-private.harmonic-reasoning-v1
Harmonic Reasoning v1
Support This Work
I'm a PhD student in visual neuroscience at the University of Toronto who also happens to spend way too much time fine-tuning, merging, and quantizing open-weight models on rented H100s and a local DGX Spark. All training compute is self-funded — balancing GPU costs against a student budget. If my open-weight models or datasets have been useful to you, consider supporting future releases.
Support on Ko-fi
Harmonic Reasoning v1 is a… See the full description on the dataset page: https://huggingface.co/datasets/DJLougen/harmonic-reasoning-v1.hsk30-graded-readers
HSK 3.0 Graded Reader Corpus
132 word-aligned Chinese graded readers with per-word pinyin and English gloss,
arranged on six difficulty shelves: 102 texts in the main split and a
disjoint 30-text held-out split.
Aligned Chinese graded-reader corpora are scarce. Existing collections are
unaligned plain text, locked inside commercial apps, or graded against HSK 2.0,
which has been superseded twice.
Loading
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/harukicoder/hsk30-graded-readers.agent-harness-paper
Thin Harness, Strong Contracts
Research artifacts for Thin Harness, Strong Contracts: Production-Oriented Agent Harnesses for Stateful AI Agents by Song Luo.
Read the full paper: Read online · PDF · Zenodo · GitHub
GitHub source: rrrrrredy/agent-harness-paper
Source commit: 586d7a37e3aa5837531cf64d5199d4ace7092ae4
Versioned research record: Zenodo DOI 10.5281/zenodo.20907471
Author: Song Luo
This Hub repository is a curated, viewer-friendly copy of the committed benchmark… See the full description on the dataset page: https://huggingface.co/datasets/RedinGhost/agent-harness-paper.repro-fixed-budget-no-harder-than-fixed-confidence-bai-traces
Agent traces
Agent sessions published from a Trackio Logbook.
web-fetch-harness-traces
Native web fetch harness traces
Separate, lightly sanitized native JSONL traces comparing URL-fetch behavior in Claude Code and Codex CLI against:
https://huggingface.co/datasets/nyu-mll/glue
Captured on 2026-09-15. No shell HTTP client, browser automation, or MCP fetcher was used.
Files
data/claude-code.jsonl: Claude Code's stream-json events.
data/codex.jsonl: Codex's persisted native rollout JSONL.
Both traces are exposed together in the default subset and… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/web-fetch-harness-traces.text-to-sql-eval-predictions
What the text-to-SQL models actually generated
Every prediction behind the numbers in
qwen3-8b-text2sql-qlora: the 453 test
questions of the enterprise text-to-SQL benchmark,
each answered by four configurations of the same model, each answer executed against the reference
PostgreSQL database and scored by comparing result sets. 1,812 rows.
I published this because the headline table (10.82 % → 50.99 % → 52.10 %) is the least interesting part of
that project. The interesting… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-eval-predictions.probe-benchmark-hard50
PROBE hard50 — human-reviewed hard development set
This is a separate standard-format review bank of 62 Astra-conditioned hard development episodes. It is NOT an independently evaluated holdout. Do not merge into or modify benchmark_600.
The original 50 human-reviewed questions were augmented with 12 human-accepted inverse comparison prompts, yielding 24 compare questions total.
Contents and ordering
Type
Count
Directories
beneath
10… See the full description on the dataset page: https://huggingface.co/datasets/vineet-datasets/probe-benchmark-hard50.text-to-sql-phrasing-robustness
Does sloppy phrasing break text-to-SQL?
The enterprise text-to-SQL benchmark
lists its own biggest caveat: every question is template-generated, so real user phrasing is untested.
This is the test. 35 test questions (one per template), each sent to the deployed
pipeline four ways: as written, with a typo, in business shorthand, and stripped to a terse fragment.
24 questions and 85 answers survive the filter described under Setup; every answer was
executed against the database.… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-phrasing-robustness.probe-benchmark-hard90
PROBE hard90
Human-reviewed hard development benchmark built from benchmark-hard50 plus accepted candidates from benchmark-hard30-review.
This bank contains 90 questions: beneath 13, compare 32, count 17, find 28. It is a hard development set, not a balanced or independent holdout.
Rejected hard30 candidates were not included. Accepted inverse compare variants were included as separate duplicate-scene questions.
CS381V-hardest-vqaasharsha30__LLAMA_Harsha_8_B_ORDP_10k-details
Dataset Card for Evaluation run of asharsha30/LLAMA_Harsha_8_B_ORDP_10k
Dataset automatically created during the evaluation run of model asharsha30/LLAMA_Harsha_8_B_ORDP_10k
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/asharsha30__LLAMA_Harsha_8_B_ORDP_10k-details.tinyperson-yolov8n-p2p3p4-hard-negative-mosaic-runsr1-h4-trigger-hardware-15fpshard-layer-v3-epistemic-honesty
VMTI Hard Layer v3: Epistemic Honesty Benchmark for Biomedical LLMs
Dataset Description
The VMTI-Trust Index (VTI) Hard Layer v3 benchmark evaluates large language models' ability to detect numerical contradictions and physiological impossibilities in clinical trial data. Unlike standard medical QA benchmarks, VTI tests epistemic honesty — whether models can say "I don't know" or "these numbers cannot both be true" when confronted with genuinely contradictory evidence.… See the full description on the dataset page: https://huggingface.co/datasets/Synho/hard-layer-v3-epistemic-honesty.hardwarize_first_ds1
Hardwarize first_ds1 (TsFile)
Apache TsFile version of Hardwarize/first_ds1.
Overview
A LeRobot dataset of keyboard-teleoperated episodes on the simulated Franka Panda
cube-pick task PandaPickCubeKeyboard-v0. Each frame records the full robot state,
the end-effector delta action, and the reinforcement-learning reward / done /
penalty signals, sampled at 50 Hz.
Task: PandaPickCubeKeyboard-v0 (Franka Panda pick-cube, keyboard teleop).
Episodes: 100 trajectories.… See the full description on the dataset page: https://huggingface.co/datasets/THULab/hardwarize_first_ds1.levir-yolov8n-p2p3p4-hard-negative-mosaic-runs
