datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
do-not-answer
Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs
Overview
Do not answer is an open-source dataset to evaluate LLMs' safety mechanism at a low cost. The dataset is curated and filtered to consist only of prompts to which responsible language models do not answer.
Besides human annotations, Do not answer also implements model-based evaluation, where a 600M fine-tuned BERT-like evaluator achieves comparable results with human and GPT-4.
Instruction… See the full description on the dataset page: https://huggingface.co/datasets/LibrAI/do-not-answer.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes.Fact-Completion
Dataset Card
Homepage: https://bit.ly/ischool-berkeley-capstone
Repository: https://github.com/daniel-furman/Capstone
Point of Contact: daniel_furman@berkeley.edu
Dataset Summary
This is the dataset for Polyglot or Not?: Measuring Multilingual Encyclopedic Knowledge Retrieval from Foundation Language Models.
Test Description
Given a factual association such as The capital of France is Paris, we determine whether a model adequately "knows" this… See the full description on the dataset page: https://huggingface.co/datasets/Polyglot-or-Not/Fact-Completion.Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.noteflow-research-pilots
Keep the failed attempts. Check the artifact.
Versioned public development evidence from Robot Reel × Skills Anywhere × EvalArc, recorded 14 September 2026 on an NVIDIA L40S, with separate scripted Harbor controls on CPU and separate GPU context-control and agent-requested MCP handoff cohorts recorded 19 September 2026. This is an inspectable engineering casebook, not a held-out benchmark or training corpus with established efficacy.
Configuration
Actual experiment
What… See the full description on the dataset page: https://huggingface.co/datasets/glayguo/noteflow-research-pilots.harvey-notes-v4
wm-rl notes v4 — two note banks from a recursive self-experience loop
Continuation update (2026-09-14): rounds 6–15 appended. The recursive bank now contains 308,580 notes / 138,723,516 training tokens. The original 108,081-row round-5 bank remains an exact prefix. Round-4/5 task lists were available for duplicate rejection, but their trajectory manifests were not published; the continuation therefore seeds round 6 from the published recall sessions and carries complete… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-notes-v4.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 5.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think.fable5-repos
Fable 5 — All-Commits GitHub Repositories
A collection of 7,090 public GitHub repositories whose entire default-branch
history was written by Claude Fable 5 — every non-merge commit carries the
trailer:
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Each repository is stored as a full .tar.gz archive including its complete
.git history, so you get every commit, message, and diff exactly as it
appears on GitHub. A manifest.jsonl / manifest.csv table describes every
repo… See the full description on the dataset page: https://huggingface.co/datasets/notune/fable5-repos.SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched
VmaxRL/SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched
This dataset contains a 350-row subset selected from the Indist SWEUniverse training rows.
Selection policy: three-way repo overlap with Bugpilot and LM-Modify, deduped by repo plus introduction patch, then balanced round-robin across overlapping repos.
Rows: 350
Selected repos: 19
Deduped overlap capacity: 468
Source dataset: VmaxRL/SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 1.3000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 4.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think.NotAllCodeIsEqual
NotAllCodeIsEqual
This dataset was created for the paper Not All Code Is Equal: A Data-Centric Study of Code Complexity and LLM Reasoning.
It contains code fine-tuning datasets split by complexity metrics for studying the relationship between code complexity and reasoning capabilities.
We provide 2 types of dataset, that cover complementary settings:
CodeNet (solution-driven complexity):
The CodeNet splits contain the same programming problems across all complexity levels, but with… See the full description on the dataset page: https://huggingface.co/datasets/itsluketwist/NotAllCodeIsEqual.NuminaMath-CoT-Small-215k
Summary
This dataset is a scaled down version of the original AI-MO/NuminaMath-CoT dataset.
Source breakdown
Source
Number of Originial Samples
Number of Samples in This Dataset
aops_forum
30201
7548
amc_aime
4072
1017
cn_k12
276591
69138
gsm8k
7345
1835
math
7478
1869
olympiads
150581
37640
orca_math
153334
38328
synthetic_amc
62111
15527
synthetic_math
167895
41968
Total
859608
214870
NuminaMath-CoT-Small-Hard-200k
Summary
This dataset is a scaled down version of the original AI-MO/NuminaMath-CoT dataset with more focus on hard math.
Source breakdown
Source
Number of Originial Samples
Number of Samples in This Dataset
aops_forum
30201
5000
amc_aime
4072
4070
cn_k12
276591
55310
gsm8k
7345
1000
math
7478
1000
olympiads
150581
37640
orca_math
153334
30662
synthetic_amc
62111
31054
synthetic_math
167895
33574
Total
859608
199310
epstein-emails
Epstein Email Threads Dataset
Dataset Summary
This dataset contains 5,082 parsed email threads extracted from OCR'd documents released by the U.S. House Oversight Committee. The emails have been processed using large language models to extract structured information including senders, recipients, timestamps, subjects, and message bodies, with OCR errors corrected and footers removed.
Dataset Description
Overview
This is a structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/notesbymuneeb/epstein-emails.mac-app-store-apps-release-notes
Dataset Card for Macappstore Applications Release Notes
📌 Dataset status: static snapshot (no scheduled updates). This dataset is derived from the December 2023 – January 2024 Mac App Store metadata snapshot and reflects the store as of that period. The dataset is stable and remains available for research use; it is not refreshed on a schedule.
Mac App Store Applications release notes extracted from the metadata from the public API.
Curated by: MacPaw Way Ltd.
Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/mac-app-store-apps-release-notes.notabug-code
NotaBug Code Dataset
Dataset Description
This dataset was compiled from code repositories hosted on NotaBug.org, a free code hosting platform that emphasizes software freedom and privacy. NotaBug is built on a fully free software stack and is popular among free software advocates and privacy-conscious developers.
Dataset Summary
Statistic
Value
Total Files
12,622,961
Total Repositories
11,660
Total Size
12 GB (compressed Parquet)
Programming… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/notabug-code.wmrl-v4-note-conditioned-rollouts-20
20 fresh note-conditioned Qwen3.5-9B actor trajectories
These are complete tool-using agent trajectories. Every task executed 2–11
document tools: 56 read, 11 grep, and 2 glob calls in total. No generated
reasoning, executed tool call, returned observation or final answer was removed.
The base actor system is byte-identical to the published base agentic eval system;
the teacher-only memory instruction and notes were appended to it.
The default table begins with tool_sequence… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/wmrl-v4-note-conditioned-rollouts-20.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 2.2000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think.Kalomaze-Opus-Instruct-25k-filteredFiltered version of Kalo's Opus_Instruct_25k for use in Celeste dataset
25.07.2024
Filter 1: removed rows which have "Claude" in responses
Filter 2: removed rows with majority of non-english text
Filter 3: Swap system prompt
Filter 4: Swap "Claude" to "Celeste" in inputs
F3 and F4 are uploaded as a separate file
amazon-c11-nothink-distillation-filtered
Amazon c11 no-think quality-filtered distillation
This preserves the selected C11 writer/criterion-judge examples and their row geometry while discarding all teacher scratch reasoning. Membership follows the pinned C11 quality-filter policy. Signed aggregate filter provenance is under quality/<split>/; it is not exposed as another dataset configuration.
The six configurations cross the two candidate variants with the three frozen stopping objectives. Every
configuration exposes… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c11-nothink-distillation-filtered.Opus-4.6-RU-Reasoning-creative-1385x-not-filtered
Opus-4.6-RU-Creative-Writing — Russian Creative Writing Reasoning Dataset
A Russian-language dataset of creative writing tasks generated with Claude claude-opus-4.6 (extended thinking enabled). Each sample contains a creative prompt, a full reasoning chain showing the creative process, and a detailed artistic response.
Dataset Info
Language: Russian 🇷🇺
Size: ~1,385 samples (growing)
Model used: anthropic/claude-opus-4.6 with reasoning: {effort: "high"}
Format:… See the full description on the dataset page: https://huggingface.co/datasets/DarkyMan/Opus-4.6-RU-Reasoning-creative-1385x-not-filtered.augmented-clinical-notesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/augmented-clinical-notes.ruwiki-pretrain-20260102ru-reasoning_effort-sft_dpo_think_gpt
NotEvilAI/ru-reasoning_effort-sft_dpo_think_gpt
NotEvilAI/ru-reasoning_effort-sft_dpo_think_gpt -
синтетический датасет для поддержки генерации ризонинга на русском языке с вариативным объёмом thinking(reasoning_effort).
Reasoning_effort представлен в виде системного промта Reasoning: [effort], где effort - одно из следующих значений:
low, medium, high - стандартные значения минимального, среднего и большого ризонинга для gpt-oss-20b/gpt-oss-120b
none - отключить ризонинг, в… See the full description on the dataset page: https://huggingface.co/datasets/NotEvilAI/ru-reasoning_effort-sft_dpo_think_gpt.amazon-c11-nothink-distillation
Amazon c11 no-think distillation
This preserves the selected C11 writer/criterion-judge examples and their row geometry while discarding all teacher scratch reasoning. Membership follows the complete pinned C11 stopping-objective corpus.
The six configurations cross the two candidate variants with the three frozen stopping objectives. Every
configuration exposes only its combined training view and preserves the native train, validation, and test
splits.
Config
Train… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c11-nothink-distillation.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/johnny8808/augmented-clinical-notes.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/Vinay393/augmented-clinical-notes.ted-polish-procurement-notices
TED Polish procurement notices
Polish narrative text rebuilt from official TED procurement-notice records.
Indexed notices: 20,588
Search API records: 20,588
XML fallbacks: 1,255
Retained documents: 15,544
Tokens: 42,475,098 (cl100k_base proxy)
Dates: 2023-02-15 to 2024-05-15
Contracting-authority attribution: 100.0%
The pinned PleIAs mirror is an identifier index only because its Polish preview contains replacement-character encoding damage. Released text comes from official… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/ted-polish-procurement-notices.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.05.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think.
