datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SHP
🚢 Stanford Human Preferences Dataset (SHP)
If you mention this dataset in a paper, please cite the paper: Understanding Dataset Difficulty with V-Usable Information (ICML 2022).
Summary
SHP is a dataset of 385K collective human preferences over responses to questions/instructions in 18 different subject areas, from cooking to legal advice.
The preferences are meant to reflect the helpfulness of one response over another, and are intended to be used for training RLHF… See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/SHP.StackMathQA
StackMathQA
StackMathQA: A Curated Collection of 2 Million Mathematical Questions and Answers Sourced from Stack Exchange
StackMathQA is a meticulously curated collection of 2 million mathematical questions and answers, sourced from various Stack Exchange sites. This repository is designed to serve as a comprehensive resource for researchers, educators, and enthusiasts in the field of mathematics and AI research.
Configs
configs:
- config_name: stackmathqa1600k… See the full description on the dataset page: https://huggingface.co/datasets/math-ai/StackMathQA.LiveSports-3K
LiveSports-3K Benchmark
News
[2025.05.12] We released the ASR transcripts for the CC track. See LiveSports-3K-CC.json for details.
Overview
LiveSports‑3K is a comprehensive benchmark for evaluating streaming video understanding capabilities of large language
and multimodal models. It consists of two evaluation tracks:
Closed Captions (CC) Track: Measures models’ ability to generate real‑time commentary aligned with the
ground‑truth ASR transcripts.
Question… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/LiveSports-3K.SHP-2
🚢 Stanford Human Preferences Dataset v2 (SHP-2)
Summary
SHP-2 is a dataset of 4.8M collective human preferences over responses to questions/instructions in 129 different subject areas, from cooking to legal advice. It is an extended version of the original 385K SHP dataset.
The preferences are meant to reflect the helpfulness of one response over another, and are intended to be used for training RLHF reward models and NLG evaluation models (e.g., SteamSHP).
Each example… See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/SHP-2.Qwen3.6-35B-A3B-mcr-stage-b
Qwen3.6-35B-A3B — MCR Stage B Corpus (Distributed Reasoning Localization)
First systematic mechanistic-intervention corpus on a hybrid MoE + GDN + Gated-Attention architecture.
📄 Paper: Loop-Intolerance Profiling: Localizing Distributed Reasoning in a Hybrid MoE Architecture via Nine Convergent Intervention Experiments — submitted to arXiv (2026-04-20, in moderation). Final arXiv ID will be added here once approved.
This dataset contains per-token residual-stream activations at… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/Qwen3.6-35B-A3B-mcr-stage-b.math-reasoning-sft-100k
Math Reasoning SFT (100K)
100,000 math problems with detailed step-by-step solutions — ready for supervised fine-tuning of math reasoning models.
Dataset Description
100,000 problems across 8 mathematical categories and 3 difficulty levels:
Categories
Category
Examples
Topics
word_problems
~23,100
Rate/time/distance, work problems, mixture, meeting/catch-up
arithmetic
~15,400
Percentages, profit/loss, ratios
geometry
~15,400
Area… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/math-reasoning-sft-100k.medical-clinical-reasoning-sft-100k
Medical Clinical Reasoning SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality medical and clinical reasoning conversations designed to train AI assistants capable of supporting clinical decision-making, documentation, and medical education.
Dataset Description
This dataset covers a broad spectrum of clinical practice scenarios across 10 medical specialty categories. Each record follows the ShareGPT conversation format with a detailed human… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/medical-clinical-reasoning-sft-100k.concurrentqa-retrievalConcurrentQA is a textual multi-hop QA benchmark to require concurrent retrieval over multiple data-distributions (i.e. Wikipedia and email data). This dataset was constructed by researchers at Stanford and FAIR, following the data collection process and schema of HotpotQA. This benchmark can be used to study generalization in retrieval as well as privacy when reasoning across multiple privacy scopes --- i.e. public Wikipedia documents and private emails.
This dataset is for the Retrieval… See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/concurrentqa-retrieval.us-statutes
US Statutes — held word for word
Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.
1,295,620 statute sections: the whole United States Code (60,433 sections, all 53 titles)
plus 27 states (1,235,187 sections), each section as its legislature publishes it, with
the URL it was… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-statutes.gpt-oss-120b-reasoning-STEM-5K
GPT-OSS-120B-Distilled-Reasoning-STEM Dataset
1) Dataset Overview
Data Source Model: gpt-oss-120b-high
Task Type: STEM Reasoning and Problem Solving (Science, Technology, Engineering & Mathematics)
Data Format: `JSON Lines
Fields: generator, category, input, CoT_Native——reasoning, answer
(Consistent with the math dataset, splitting the original 'output' into 'reasoning' and 'answer' for COT/SFT scenarios.)
2) Design Goals (Motivation)
This dataset targets… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120b-reasoning-STEM-5K.StackPulse_778K_QnA_Code_dataset
💻 StackOverflow-778K: Multi-Year Developer Q&A Dataset
Dataset Summary
A large-scale Stack Overflow question dataset containing 778,929 unique
questions sampled across 7 years (2015–2022). Each question includes the
raw HTML body, plain-text version, tags, score, view count, answer count, and
a rich set of derived features for immediate ML use.
Collected across 8 sampling runs on Feb 27 2026, deduplicated to
778,929 unique questions with only 2 duplicates removed.… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/StackPulse_778K_QnA_Code_dataset.science-qa-sft-100k
Science QA SFT (100K)
100,000 science Q&A examples with step-by-step explanations for SFT fine-tuning. Covers physics, chemistry, biology, astronomy, and earth science at beginner through advanced difficulty.
Motivation
Models trained on general text often give superficially plausible but mechanistically wrong answers to science questions — stating the right conclusion without understanding the underlying reasoning. This dataset trains models to explain why an… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/science-qa-sft-100k.salabs-stem-deep-reasoning-cot-v13
🧪 SALabs Multi-Domain STEM Deep Reasoning & Chain-of-Thought (CoT) Corpus (v13.0)
[!IMPORTANT]
💳 Click Here to Purchase Enterprise Commercial License ($2,500 USD) & Instant 31.7MB Master Archive DownloadInstant download of the full lossless master package containing all 1,816 JSONL reasoning records + 13 complete uncompressed text corpora (31.72 MB uncompressed total) + commercial license certificate.
🌟 Executive Summary
The SALabs STEM Deep Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/suitai/salabs-stem-deep-reasoning-cot-v13.StackMathQA
StackMathQA
StackMathQA is a meticulously curated collection of 2 million mathematical questions and answers, sourced from various Stack Exchange sites. This repository is designed to serve as a comprehensive resource for researchers, educators, and enthusiasts in the field of mathematics and AI research.
Configs
configs:
- config_name: stackmathqa1600k
data_files: data/stackmathqa1600k/all.jsonl
default: true
- config_name: stackmathqa800k
data_files:… See the full description on the dataset page: https://huggingface.co/datasets/agicorp/StackMathQA.Electrical-engineering
To the electrical engineering community
This dataset contains Q&A prompts about electrical engineering, Kicad's EDA software features and scripting console Python codes.
Authors
STEM.AI: stem.ai.mtl@gmail.comWilliam Harbec
policystrategies-archive
🏛️ Open-Source Macro-Strategy, Financial History & Intelligence Archive
🌐 Overview & Institutional Mission
This public repository serves as the official open-source knowledge graph and metadata registry for r/policystrategies.
We aggregate, document, and cross-reference declassified historical intelligence dossiers, sovereign debt crises, systemic market manipulations, and geoeconomic conflicts using verified open-source intelligence (OSINT) and primary… See the full description on the dataset page: https://huggingface.co/datasets/stratigahq/policystrategies-archive.mm-long-storytelling-bench
MM Long Storytelling Bench — v3
⚠️ The 756 model-drafted questions have been WITHDRAWN from this
dataset's splits (2026-08-05). They were drafted by a model that is also
an evaluation target, which makes them circular as a measurement
instrument. They are kept in full, with the reasoning, under
data/v3/archive/ — nothing was deleted.
The splits currently hold 6 worked examples (status: "example"),
which document the required format and are not a benchmark. Do not use
this… See the full description on the dataset page: https://huggingface.co/datasets/luoojason/mm-long-storytelling-bench.Stable-Code-Python-SFT
Stable Code Python SFT
The Stable Code Python SFT dataset is a high-quality synthetic dataset derived from the
stabilityai/stable-code-instruct-3b model for the purpose of supervised fine-tuning (SFT). Please refer to the
Versioning section for dataset versions.
Note: If you would like to contribute to this repository,
please read the CONTRIBUTING first.
TableofContents
Features
File Structure
Metadata
Usage
Versioning
License
TeamContact
Reference
Citation… See the full description on the dataset page: https://huggingface.co/datasets/bunyaminergen/Stable-Code-Python-SFT.PRAGMA
🧩 PRAGMA
PRAGMA is a benchmark for evaluating personalized guidance with memory alignment over lifelong conversation histories. It tests whether a model can use relevant past interactions while remaining aligned with a user's evolving experiences and trajectory.
PRAGMA accompanies the paper “PRAGMA: Evaluating Personalized Guidance with Memory Alignment in Lifelong Conversations” (EMNLP 2026).
✨ Dataset Overview
Statistic
Value
Users
100
Queries
400… See the full description on the dataset page: https://huggingface.co/datasets/stellahj/PRAGMA.Tool-Star-SFT-54KThe cold start dataset of Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
hf paper link: https://huggingface.co/papers/2505.16410
arxiv link: https://arxiv.org/abs/2505.16410
Github Repo: https://github.com/dongguanting/Tool-Star
cairo-security-audits
Cairo Security Audits
A source-traceable corpus of public Cairo and Starknet security-audit metadata and normalized finding annotations.
Version 0.3.0 packages every entry in the audit inventory frozen at keep-starknet-strange/starknet-skills@17a76e8. It covers 32 accessible reports from 10 auditing firms and 286 normalized finding annotations. Eleven records are checked against rendered reports and two link to exact vulnerable/fixed commits. The release does not redistribute… See the full description on the dataset page: https://huggingface.co/datasets/starknet-ai/cairo-security-audits.delvantic-stock-knowledge-layer
Delvantic Stock Knowledge Layer
A 872k-word, source-cited textbook of stock analysis and trading, organized as a tree —
the reference layer behind a live AI research engine, published in full.
Every finance dataset on the Hub is numbers: prices, filings, labelled headlines. This is the
missing other half — the explanations. 771 documents on how the machinery of markets
actually works, from reading a cash-flow statement to why volatility regimes break strategies,
each one written… See the full description on the dataset page: https://huggingface.co/datasets/fatcat55/delvantic-stock-knowledge-layer.StreetVision-10K
StreetVision-10K
Each sample contains:
A system prompt instructing the model to act as an OSINT/geospatial expert
A user message with a street-level photo and the instruction to determine coordinates
An assistant response with ground-truth coordinates in <direct_lon_lat_output>longitude,latitude</direct_lon_lat_output> format
Format
Each line is a JSON array of ChatML messages:
[
{"role": "system", "content": "..."},
{"role": "user", "content": [… See the full description on the dataset page: https://huggingface.co/datasets/mishl/StreetVision-10K.sft-tool-calling-structured-output-v1
vericava/sft-tool-calling-structured-output-v1
Dataset to train (SFT) 3-20B LLMs for tool calling and structured outputs/classifications.
Includes contents in English as well as some Japanese.
Law-StackExchange
Law-StackExchange Dataset Details
All StackExchange legal questions and their answers from the Law site, up to 14 August 2023.
The repository includes a notebook for the process using the official StackExchange API.
Citation
@misc{Moslem2023-LawStackExchangeDataset,
author = {Moslem, Yasmin},
title = {Law-StackExchange Dataset},
year = 2023,
url = {https://huggingface.co/datasets/ymoslem/Law-StackExchange},
doi =… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/Law-StackExchange.studybench
studybench
studybench is a small, high-effort benchmark of expert-level coding questions about real
open-source codebases, each paired with a gold answer and a weighted, source-grounded grading
rubric. The questions ask a model to produce working code that uses a specific library/framework
correctly; the rubric decomposes a correct answer into discrete, checkable claims, each tied to exact
lines of the upstream source.
This release publishes both the questions and the full… See the full description on the dataset page: https://huggingface.co/datasets/jacobli/studybench.StethoBench
StethoBench
StethoBench is a comprehensive benchmark for cardiopulmonary auscultation, comprising 77,027 instruction–response pairs synthesized from 16,125 labeled recordings across 11 public datasets. It is the training and evaluation benchmark for StethoLM, published in the Transactions on Machine Learning Research (TMLR).
Dataset Description
StethoBench was constructed by synthesizing instruction–response pairs from existing labeled cardiopulmonary audio datasets… See the full description on the dataset page: https://huggingface.co/datasets/askyishan/StethoBench.bespoke-stratos-es
Bespoke-Stratos-ES
Spanish reasoning traces regenerated from bespokelabs/Bespoke-Stratos-17k -- 16709 rows, natively generated in Spanish (not machine-translated from the English traces).
Models trained on this dataset
axiom-of-choice/qwen3-4b-es-reasoning-qlora (Qwen3-4B) -- also for transformers/peft: qwen3-4b-es-reasoning-peft
axiom-of-choice/qwen3-1.7b-es-reasoning-lora (Qwen3-1.7B) -- also: qwen3-1.7b-es-reasoning-peft… See the full description on the dataset page: https://huggingface.co/datasets/axiom-of-choice/bespoke-stratos-es.k12-standards-instruction-tasks
K-12 Curriculum Tasks (generated)
2,489 generated instruction/input/output records covering five curriculum tasks:
assessment creation, learning objective generation, misconception detection, standard
explanation, and standards Q&A. Content is predominantly mathematics.
Important: the name is misleading
Despite the name, this dataset contains no school directory data. There are four
columns - task, input, output, metadata - and no staff, principal, or school… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-standards-instruction-tasks.EverMemBench-Static
EverMemBench-S: Evaluating Evidence Access under Dense Semantic Interference
💻 Code: EverMind-AI/EverMemBench-Static
Overview
EverMemBench-S (EMB-S) is an adversarial Needle-in-a-Haystack benchmark built on a 326M-token MemoryBank with 160,280 documents across 8 domains. It evaluates long-context models and retrieval systems under dense semantic interference — where near-miss documents create realistic confusion that standard NIAH benchmarks cannot capture.
1,225… See the full description on the dataset page: https://huggingface.co/datasets/EverMind-AI/EverMemBench-Static.
