datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Traditional-Chinese-Common-Crawl-NOT-CleanedCommon Crawl Dumps that were briefly filtered by keywords to remove bad words and simplified Chinese.
The hash based cleaned dataset can be found here.
Files here are for future usage (downloading from Common Crawl and keyword filtering are very slow)
takedown-notices
Takedown notices received by the Hugging Face team
Please click on Files and versions to browse them
Also check out our:
Terms of Service
Community Code of Conduct
Content Guidelines
do-not-answer
Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs
Overview
Do not answer is an open-source dataset to evaluate LLMs' safety mechanism at a low cost. The dataset is curated and filtered to consist only of prompts to which responsible language models do not answer.
Besides human annotations, Do not answer also implements model-based evaluation, where a 600M fine-tuned BERT-like evaluator achieves comparable results with human and GPT-4.
Instruction… See the full description on the dataset page: https://huggingface.co/datasets/LibrAI/do-not-answer.math-contests-2026
Math Contests 2026 (🔗 notadib/math-contests-2026)
197 problems from national olympiads and team-selection tests held January 2026 and onward — a held-out benchmark for math reasoning, sourced after the contests ran but before solutions were widely propagated, so they should not appear in any current LLM training data.
Excluded: any contest held in 2025 — BMO Round 1 (Nov 2025), USA TSTST, USA TST (Dec 2025) and Bundeswettbewerb Mathematik (Dec 2025) — kept strictly to events… See the full description on the dataset page: https://huggingface.co/datasets/notadib/math-contests-2026.truthfulness_high_quality
Dataset Card for "truthfulness_high_quality"
More Information needed
Notebook_ScriptscdnSWE-bench_Not_Verifiedck12-tqa-instruction
CK-12 TQA: Textbook Question Answering (Instruction Format)
Dataset Description
Dataset Summary
This is a reformatted version of the TQA (Textbook Question Answering) dataset, converted into an instruction-following format suitable for training and evaluating large language models on science question answering and multimodal reasoning tasks.
The TQA dataset consists of 1,076 lessons from Life Science, Earth Science, and Physical Science textbooks sourced from… See the full description on the dataset page: https://huggingface.co/datasets/notefill/ck12-tqa-instruction.misconceptions_tf
Dataset Card for "misconceptions_tf"
More Information needed
NotInject
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
Website, Paper, Code, Demo
Dataset Description
The NotInject is a benchmark designed to evaluate the extent of over-defense in existing prompt guard models against prompt injection. All samples in the dataset are benign but contain trigger words that may be mistakenly flagged as risky. The dataset is divided into three subsets, each consisting of prompts generated using one… See the full description on the dataset page: https://huggingface.co/datasets/leolee99/NotInject.nota
Dataset Card for Nota
Dataset Summary
This data was created by the public institution Nota, which is part of the Danish Ministry of Culture. Nota has a library audiobooks and audiomagazines for people with reading or sight disabilities. Nota also produces a number of audiobooks and audiomagazines themselves.
The dataset consists of audio and associated transcriptions from Nota's audiomagazines "Inspiration" and "Radio/TV". All files related to one reading of one edition… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/nota.truthfulness_all
Dataset Card for "truthfulness_all"
More Information needed
Fact-Completion
Dataset Card
Homepage: https://bit.ly/ischool-berkeley-capstone
Repository: https://github.com/daniel-furman/Capstone
Point of Contact: daniel_furman@berkeley.edu
Dataset Summary
This is the dataset for Polyglot or Not?: Measuring Multilingual Encyclopedic Knowledge Retrieval from Foundation Language Models.
Test Description
Given a factual association such as The capital of France is Paris, we determine whether a model adequately "knows" this… See the full description on the dataset page: https://huggingface.co/datasets/Polyglot-or-Not/Fact-Completion.easyr1-grounding-dataset-30k-not_grounded-SE-GUI-3B-2MPbo_or_not
Dataset Card for bo-dataset
This is a FiftyOne dataset with 169 samples designed for binary classification of Bo (Barack Obama's Portuguese Water Dog) versus other pets.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/bo_or_not")
#… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/bo_or_not.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes.us-warn-act-layoffs-notices-daily
US WARN Act Layoff Notices — normalized, 48 states, rebuilt every day
Last rebuilt: 2026-09-22. An automated pipeline re-scrapes 48
state labor-department portals every day, re-normalizes, re-deduplicates and
re-uploads this file. Compare that date with the "last modified" date on any
other US WARN dataset on the Hub before you choose one — WARN data is a
moving target and a one-shot upload starts rotting the week it is posted
(states amend headcounts, re-issue notices, and… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-warn-act-layoffs-notices-daily.design-patents-not-in-impact
US Design Patents Not Included in IMPACT (2008-2026)
Original drawing images (TIFF) and grant full-text XML for 165,917 US design patents that are
absent from the AI4Patents/IMPACT dataset.
IMPACT covers 2007-2022 and contains 434,498 rows. This dataset supplies the design patents that
IMPACT does not have: 161,093 patents granted in 2023-2026, which are outside IMPACT's period,
plus 4,824 patents from years IMPACT does cover but did not include. There is no patent
overlap with… See the full description on the dataset page: https://huggingface.co/datasets/SoichiOnozuka/design-patents-not-in-impact.issues-kaggle-notebooks
GitHub Issues & Kaggle Notebooks
Description
GitHub Issues & Kaggle Notebooks is a collection of two code datasets intended for language models training, they are sourced from GitHub issues and notebooks in Kaggle platform. These datasets are a modified part of the StarCoder2 model training corpus, precisely the bigcode/StarCoder2-Extras dataset. We reformat the samples to remove StarCoder2's special tokens and use natural text to delimit comments in issues and display… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/issues-kaggle-notebooks.notch-beam-2d-impact
NotchBeam2D-Impact — StructBench canonical dataset
Download
One case, one file — fetch exactly what you need (pip install huggingface_hub):
from huggingface_hub import hf_hub_download, snapshot_download
# one case
path = hf_hub_download("StructBench/notch-beam-2d-impact",
filename="<case_id>.h5", repo_type="dataset")
# the full archive (resumable; cached under HF_HOME)
root = snapshot_download("StructBench/notch-beam-2d-impact"… See the full description on the dataset page: https://huggingface.co/datasets/StructBench/notch-beam-2d-impact.Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.spider_mcqa_v0.2_full
Spider-MCQA
Converted Spider Text-to-SQL (Paper: Yu et al., 2018; HF Dataset) test set into multiple-choice.
The dataset contains 1,034 examples.
Dataset Fields
Each JSON record contains:
query: the schema and natural-language question prompt.
gold_answer: the correct SQL answer.
options: four SQL answer options, including the gold answer and three generated distractors.
correct_option_index: the index of the correct answer in options.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/notpaulmartin/spider_mcqa_v0.2_full.insider-sale-notices
US Insider Sale Notices — SEC Form 144
Every notice a corporate insider files before selling restricted or control
stock, normalized into a point-in-time schema and rebuilt daily.
124 886 notices covering 2023-01-04 to 2026-09-04, from
4 237 issuers and 23 871 sellers, with the acceptance timestamp
of every filing to the second.
The pipeline that produces this dataset lives in recipe/ inside
this same repository, at the same revision as the data. See
PIPELINE.md for the method… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/insider-sale-notices.azaria-mitchell
Dataset Card for "azaria-mitchell"
More Information needed
python_functions_reasoningThis is the Python (functions) coding reasoning dataset used to train
Notbad v1.0 Mistral 24B reasoning model.
The reasoning data were sampled from an RL-based self-improved
Mistral-Small-24B-Instruct-2501 model.
The Python functions and instructions were sourced from OpenCoder Dataset Stage1
and from open source projects on Github.
You can try Notbad v1.0 Mistral 24B on chat.labml.ai.
ACE-Trajectories_noTossesPaper in the making
ACE-Trajectories_noTosses Dataset
This dataset was created for the Master's thesis "From Broadcast to 3D: A Deep Learning Approach for Tennis Trajectory and Spin Estimation" by Alexandra Göppert at the University Augsburg, Chair of Machine Learning and Computer Vision.
This datasets serves as an enriched version of the original TrackNet Tennis dataset. It cuts the whole rallies included in the TrackNet dataset in single trajectories, based on TrackNets property… See the full description on the dataset page: https://huggingface.co/datasets/XSpaceCoderX/ACE-Trajectories_noTosses.us-warn-act-layoff-notice-period-days
How much notice did US layoff notices actually give?
The WARN Act is, at bottom, a law about a number of days. Every state
publishes layoff notices; none of them publishes the one column that says
whether the notice arrived in time. This dataset is that column, recomputed
every day from the primary filings.
notice_days = effective_date - notice_date, per notice.
The headline, over the full archive (61,330 notices, 1988 - today)
Notices stating both a… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-warn-act-layoff-notice-period-days.ivirits-audio-v2-30s
ivrit.ai audio-v2 — 2–30 s segments
ivrit-ai/audio-v2 (>20k hours of Hebrew
audio) cut into 2–30 second speech segments with machine transcripts, ready for ASR
fine-tuning.
How it was built
VAD — Silero VAD (ONNX) over each episode decoded to 16 kHz mono. Speech regions
longer than 30 s are split at the quietest sufficiently-long pause inside the window,
so cuts land in silence rather than mid-word. Regions shorter than 2 s are dropped.
Transcription —… See the full description on the dataset page: https://huggingface.co/datasets/notmax123/ivirits-audio-v2-30s.us-layoffs-by-stock-ticker-warn-act-notices-public-companies
US layoffs by stock ticker: 7,651 WARN Act notices filed by 736 listed companies, with confidence tiers
Rebuilt 2026-09-22. State WARN Act filings name the employer as the filer wrote
it — Wells Fargo Home Mortgage, OS Restaurant Services, Boeing Compnay —
never a ticker. This dataset resolves those strings to the listed parent and
re-derives the mapping every day as new filer strings appear.
1,583 filer strings → 736 tickers → 7,651 notices
(1988-12-16 → 2026-09-16), 1,007,332… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-layoffs-by-stock-ticker-warn-act-notices-public-companies.
