datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kernelbench-hard-traces
KernelBench-Hard agent traces
Frontier coding agents writing optimized CUDA/Triton kernels (FP8 GEMM, paged
attention, MoE, W4A16, KDA, Top-k) on RTX PRO 6000 Blackwell, H100 PCIe, and
B200; roofline-graded.
Each .jsonl file is one agent run in Claude-Code session format, viewable with
the Hugging Face Agent Trace viewer (Data Studio → open a row). Filename =
run id.
Live leaderboard: https://kernelbench.com/hard
Secrets redacted. Full reasoning for open-provider routes… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-traces.newswire
Dataset Card for NewsWire
Dataset Summary
NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model.
Languages
English (en)
Dataset Structure
Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.HarmfulQAPaper | Github | Dataset| Model
📣📣📣: Do check our new multilingual dataset CatQA here used in Safety Vectors:📣📣📣
As a part of our research efforts toward making LLMs more safe for public use, we create HarmfulQA i.e. a ChatGPT-distilled dataset constructed using the Chain of Utterances (CoU) prompt. More details are in our paper Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment
HarmfulQA serves as both-a new LLM safety benchmark and an alignment dataset… See the full description on the dataset page: https://huggingface.co/datasets/declare-lab/HarmfulQA.terminal-bench-2-verified
Terminal-Bench 2.0 Verified: Instruction & Environment Fix Version
中文版本
We conducted a comprehensive review of the entire Terminal-Bench 2.0 dataset and identified various issues. Both GLM-5 and Step 3.5-Flash were evaluated using this verified version.
This modified version addresses environment and instruction issues we discovered in Terminal-Bench 2.0. It includes two types of fixes:
Environment Fixes: Updated Dockerfiles and instructions to support Claude Code Agent runtime… See the full description on the dataset page: https://huggingface.co/datasets/harithoppil/terminal-bench-2-verified.harbor-swesmith-rl-artifacts
Harbor SWE-Smith 强化学习数据产物
本数据集是 Harbor Qwen 工具调用代码智能体强化学习项目使用的冻结任务集,服务于 GRPO、原生价值模型/GAE PPO、训练过程诊断和统一协议评测。
项目已于 2026 年 8 月 30 日完成 P0 评测并进入阶段性归档。本数据集用于保留实验所依赖的数据切分、任务执行文件和审计信息,不代表新的通用代码能力基准。
数据概况
切分
任务数
训练集
187
验证集
42
测试集
38
合计
267
数据覆盖 89 个上游代码仓库。三个切分之间同时执行任务标识和仓库级隔离检查。
正式数据集名称:
swesmith-curated-grpo-267-v1
冻结切分的语义摘要:
ae5df9a3f4a3fc8af44fac420b36529e283839e1bd3de9daba65d5bcda51447d
该值来自 split-manifest.json 的 sha256 字段,用于标识切分语义,不等同于该文件本身的字节级… See the full description on the dataset page: https://huggingface.co/datasets/keryszhan/harbor-swesmith-rl-artifacts.DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
DeepSeek v4 Pro Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-pro.
JSONL files: 4006
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/hardcoremoore/DeepSeek-v4-Pro-Agent.volume2gym-railroad-1959
The Rulebook Becomes a World
Volume2Gym asks a deliberately expansive question: what if any sufficiently structured text could become a small, inspectable world in which a model learns by acting, receiving feedback, and trying again?
This release turns one bounded English technical volume into an auditable reinforcement-learning dataset: 117 source scans → 536 extracted rules → 2,708 synthetic scenario tasks → a measured rule-linkage and verification surface. It is an… See the full description on the dataset page: https://huggingface.co/datasets/HarleyCooper/volume2gym-railroad-1959.gemini-3.1-pro-hard-high-reasoning
Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M
Dataset Details
Dataset Description
This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification.
The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3.1-pro-hard-high-reasoning.MiMo-2.5-Pro-Reasoning-Traces-Hard
MiMo-2.5-Pro-Reasoning-Traces-Hard
A large-scale reasoning dataset of 8,706 expert-level prompts with full reasoning traces across 44 academic and technical topics, generated using the MiMo-v2.5-Pro model. Each entry contains the step-by-step reasoning chain alongside the final completion, designed for training and evaluating advanced reasoning capabilities in language models.
Dataset Statistics
Metric
Value
Total entries
8,706
Unique topics
44… See the full description on the dataset page: https://huggingface.co/datasets/Skyhigh-2203/MiMo-2.5-Pro-Reasoning-Traces-Hard.HardGen
From Failure to Mastery: Generating Hard Samples for Tool-use Agents
This is only a demonstration result obtained by directly deploying and sampling HardGen in the BFCL environment, not the dataset directly used in the HardGen paper.
[!IMPORTANT] Important Hint
This is an extension of the technical report FunReason-MT Technical Report: Advanced Data Synthesis Solution for Real-world Multi-Turn Tool-use
To allow the model to learn from errors, we specifically construct… See the full description on the dataset page: https://huggingface.co/datasets/Bingguang/HardGen.terminal-bench-2-trajectories
Terminal-Bench 2.0 Leaderboard Trajectories
Agent trajectories extracted from Terminal-Bench 2.0 leaderboard submissions. Each row contains a prompt (task instruction), the agent's response, and the reward (pass/fail).
Models Included
Model
Trials
Passed
Claude-Opus-4.6
2,213
1,537 (69%)
Gemini-3.1-Pro-Preview
445
333 (75%)
GLM-5
445
231 (52%)
Kimi-k2.5
442
189 (43%)
Claude-Opus-4.5
178
98 (55%)
Splits
Config
Description
Rows… See the full description on the dataset page: https://huggingface.co/datasets/harithoppil/terminal-bench-2-trajectories.soc-audit-11k
SOC Audit Text Generation Dataset
Description
This dataset is designed for training and evaluating Language Models (LLMs) specifically in the context of SOC 2 audits. It covers a wide range of topics including, but not limited to, information security, risk management, compliance, data privacy, and governance. The dataset consists of structured text in the format of instructions followed by a detailed response, making it ideal for models intended to assist in… See the full description on the dataset page: https://huggingface.co/datasets/harleygilpin/soc-audit-11k.instruction-following-hard-sft-100k
Hard Instruction Following SFT (100K)
100,000 ShareGPT conversations where the assistant correctly satisfies multiple simultaneous explicit constraints in a single response. Each example pairs a multi-constraint prompt with a response that honors every constraint without dropping any.
Targets the instruction-following capability measured by IFEval and similar benchmarks.
Motivation
A key failure mode in deployed LLMs is dropping constraints under load — responding… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/instruction-following-hard-sft-100k.openlegaldataClean Open Legal Data
Overview |
Dataset Structure |
Key Fields |
Example Entry |
Using the Dataset with Python |
Citation |
License
Overview
This dataset is a comprehensive collection of open legal case records in JSONL format. It comprises 423,941 cases extracted and processed from the Open Legal Data dump dump-20260520 and represents an independent, cleaned derivative of that source data. The dataset is designed… See the full description on the dataset page: https://huggingface.co/datasets/harshildarji/openlegaldata.enterprise-text-to-sql-benchmark
Enterprise Text-to-SQL Benchmark
3,087 natural-language questions paired with executable PostgreSQL, over a
12-table enterprise schema (sales, catalogue, logistics, HR).
Built to answer one question honestly: does fine-tuning actually improve
text-to-SQL? On this benchmark, a QLoRA fine-tune of Qwen3-8B took strict
execution accuracy from 43.71 % to 68.43 %, and 70.86 % with a
self-correction loop — and the benchmark is designed so that number cannot be
inflated by leakage or by… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/enterprise-text-to-sql-benchmark.HarmfulSkillBench
📝 Paper |
📑 arXiv |
💻 Code |
📦 Dataset
HarmfulSkillBench
A benchmark for evaluating LLM refusal behavior when agents are exposed to skills
that describe potentially harmful capabilities.
The benchmark probes whether current LLMs can detect and refuse harmful agent
skills in two settings. Tier 1 covers prohibited behaviors that should always
be refused. Tier 2 covers high-risk domains where responses should include
human-in-the-loop referral and AI… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/HarmfulSkillBench.needle2-harness-dispatch
Needle-2 Harness-Dispatch Corpus (review build)
Eval/training corpus for tool-dispatch on a developer-agent harness surface
(8 tools: bash / read / write / edit / glob / grep / web_search /
todo_write). This is a review build — every record carries QA
annotations so a human can approve, relabel, or flag before the next
training run.
Provenance
Generated and judged by glm-5.3 , two
generation rounds (seeds 7 and 101), judge pass kept/fixed/dropped.
1,307 raw… See the full description on the dataset page: https://huggingface.co/datasets/ebowwa/needle2-harness-dispatch.ArabPref
ArabPref Preference Test
This repository contains the English and Arabic preference test data from ArabPref.
It also contains the English and Arabic MCQ test data.
Files
pref_test_en.jsonl: 3,300 English examples.
pref_test_ar.jsonl: 2,964 Arabic examples.
mcq_test_en.jsonl: 992 English multiple-choice questions.
mcq_test_ar.jsonl: 992 Arabic multiple-choice questions, including the revised items.
The two files are exposed as separate configurations because they… See the full description on the dataset page: https://huggingface.co/datasets/Skywalker-Harrison-mbz/ArabPref.biden-harris-redteam-archived
THIS IS AN ARCHIVED VERSION
Biden-Harris Redteam: A red-teaming dataset focusing on the Biden-Harris AI Executive Order
Dataset Description
While building Large Language Models (LLMs), it is crucial to protect them against attacks that could bypass safety guardrails and break their guiding principles. Specifically, LLMs should never generate content promoting or normalizing harmful, illegal, or unethical behavior that may contribute to the harm of the… See the full description on the dataset page: https://huggingface.co/datasets/aurora-m/biden-harris-redteam-archived.harmonic-reasoning-v1
Harmonic Reasoning v1
Support This Work
I'm a PhD student in visual neuroscience at the University of Toronto who also happens to spend way too much time fine-tuning, merging, and quantizing open-weight models on rented H100s and a local DGX Spark. All training compute is self-funded — balancing GPU costs against a student budget. If my open-weight models or datasets have been useful to you, consider supporting future releases.
Support on Ko-fi
Harmonic Reasoning v1 is a… See the full description on the dataset page: https://huggingface.co/datasets/DJLougen/harmonic-reasoning-v1.math-sft-solutions-no-cot
Math SFT Solutions No CoT
A cleaned mathematics supervised fine-tuning dataset containing:
instruction → solution pairs
mathematical proofs
derivations
olympiad-style solutions
theorem reasoning
stepwise mathematical explanations
detailed final solutions
This dataset was built specifically for mathematical supervised fine-tuning (SFT).
Unlike many reasoning datasets, this release removes explicit chain-of-thought tags and hidden thinking traces while preserving high-quality… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/math-sft-solutions-no-cot.hsk30-graded-readers
HSK 3.0 Graded Reader Corpus
132 word-aligned Chinese graded readers with per-word pinyin and English gloss,
arranged on six difficulty shelves: 102 texts in the main split and a
disjoint 30-text held-out split.
Aligned Chinese graded-reader corpora are scarce. Existing collections are
unaligned plain text, locked inside commercial apps, or graded against HSK 2.0,
which has been superseded twice.
Loading
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/harukicoder/hsk30-graded-readers.gemini-3-pro-10000x-hard-high-reasoning
Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning
Dataset Details
Dataset Description
Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement.
This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3-pro-10000x-hard-high-reasoning.pmpp-hard
PMPP-Hard Agent Evaluation Traces
PMPP-Hard is a 69-task agentic GPU-kernel evaluation for testing whether autonomous coding agents can produce implementations that are both correct and performant. This dataset contains the complete nine-model campaign used in the PMPP-Hard release: 621 rollouts, with 69 task sessions for each model configuration.
Maintained and released by Sinatras.
Source repository: SinatrasC/pmpp-hard
Prime environment and evaluations: PMPP-Hard on Prime… See the full description on the dataset page: https://huggingface.co/datasets/sinatras/pmpp-hard.structured-hard-sft-4k
Hard Synthetic Dataset for Structured Data Tasks (v1)
This dataset contains 4,000 high-difficulty synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks.
The data is fully synthetic, generated using deterministic serialization to ensure syntax validity while maintaining high structural complexity (deep nesting and varied types).
Dataset Summary
The dataset addresses four "hard" areas typically… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structured-hard-sft-4k.Indian-legal-data-v3
Indian Legal Dataset V3
Overview
Indian Legal Dataset V3 is a large-scale instruction-tuning dataset focused on Indian law, constitutional law, criminal law, legal reasoning, legal drafting, and real-world legal assistance.
Compared to V2, this version expands the dataset with:
legal drafting instruction pairs,
hypothetical legal scenarios,
detailed IPC-focused data,
practical real-world legal instructions,
concise legal QA pairs.
After integrating the new data sources… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Indian-legal-data-v3.math-sft-solutions-no-cot-v3
Math SFT Solutions No CoT V3
Math SFT Solutions No CoT V3 is a large-scale mathematics supervised fine-tuning (SFT) dataset designed for instruction tuning and mathematical capability adaptation.
Version 3 substantially expands mathematical coverage while improving dataset quality through stronger filtering, cleaning, and supervision refinement.
Unlike reasoning-heavy datasets, this release focuses on clean instruction → response pairs without hidden chain-of-thought style… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/math-sft-solutions-no-cot-v3.Uncensored-SFT-v1
Dataset Creation Process
This dataset was not scraped from a single source.
Instead, it was built through a large multi-stage curation and cleaning pipeline involving many open instruction datasets available on Hugging Face.
The entire dataset was normalized into a unified:
{
"input": "...",
"output": "..."
}
format.
Data Collection
A large number of public instruction datasets were downloaded from Hugging Face.
These datasets included:
Instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Uncensored-SFT-v1.IntentRouterTrain
MODF-SIR: a Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning
MODF-SIR is a lightweight MLLM-based, distillation-augmented, multi-agent collaborative framework for social intelligence reasoning.
🔖 Model Details
Model type: Omni-modal Large Language Model
License: BSD-3-Clause
Project Page: Arxiv: https://arxiv.org/abs/2606.12018
Code Repository: GitHub: https://github.com/eeee-sys/MODF-SIR
👀 MODF-SIR… See the full description on the dataset page: https://huggingface.co/datasets/Harry-1234/IntentRouterTrain.hardware-verilogeval-v2
hardware-verilogeval-v2
VerilogEval v2 - 471 Verilog evaluation problems
Dataset Overview
This dataset is part of a comprehensive collection of hardware design datasets for training and evaluating LLMs on Verilog/SystemVerilog code generation and hardware design tasks.
Files
verilog_eval_problems.json: 471 VerilogEval v2 problems
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset('AbiralArch/hardware-verilogeval-v2')… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/hardware-verilogeval-v2.
