datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sea-syntheticunclickbait-synthetic-27b-trajectories
Unclickbait Synthetic 27B Trajectories
Synthetic trajectory dataset generated by using rich structured JSON prompts and validated by two-stage judging pipeline.
Contents
: Full generated trajectories (current snapshot: 48,623 records out of 152,369 pristine event candidates).
: 30 benchmark test samples audited end-to-end through the 122B two-stage judge (Stage 1 integrity gate + Stage 2 4D scoring).
machinelearninglm-scm-synthetic-tabularml
MachineLearningLM Pretraining Corpus
This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.synthetic-pii-function-calling
Dataset Summary
A function calling dataset created by filtering the urchade/synthetic-pii-ner-mistral-v1 dataset.
pi-synthetic
Coding agent session traces for aaaaliou/pi-synthetic
This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-synthetic.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-synthetic.Synthetic-AI-ML-Dataset
Synthetic-AI-ML-Dataset
Synthetic Q&A dataset on AI and Machine Learning
Dataset Details
Metric
Value
Topic
AI and Machine Learning
Total Q&A Pairs
14021
Valid Pairs
14021
Provider/Model
ollama/gpt-oss:120b
Generation Cost
Metric
Value
Prompt Tokens
14,941,957
Completion Tokens
17,159,263
Total Tokens
32,101,220
GPU Energy
12.9628 kWh
Sources
This dataset was generated from 474 scholarly papers:
#… See the full description on the dataset page: https://huggingface.co/datasets/Kylan12/Synthetic-AI-ML-Dataset.Synthetic-Causal-Reasoning-50k
🏭 Sovereign Synthetic Reasoning Dataset (400k)
"High-Quality Chain-of-Thought Data at Scale."
📊 Overview
This dataset contains 400,000 synthetic reasoning samples spanning 16 enterprise domains (Finance, Pharma, Legal, Cybersecurity, Supply Chain, etc.).
It was generated using the Sovereign Generator, which produced 1.6 million samples and applied a strict quality filter (Top 25%) to retain only the most logically consistent and complex chains.
Average Quality… See the full description on the dataset page: https://huggingface.co/datasets/davidfoss/Synthetic-Causal-Reasoning-50k.synthetic-beir-dataBILGE-Synthetic-Web
BILGE-Synthetic-Web Dataset
BILGE-Synthetic-Web was created following the methodology presented in the Cosmopedia blog/article.
All content was generated using a 27B-parameter model.
Further details on the methodology are available at:
🔗 https://huggingface.co/blog/cosmopedia
synthetic-social-networks
Synthetic Social Networks (Dataset)
Raw experimental outputs from the Synthetic Social Networks study:
59,776 in-character LLM-agent posts from 528 production trials, and
64,562 posts total when the original pipeline-verification runs are
included. The artifact combines an exploratory stage with a separately
frozen, preregistered 448-trial matched-exposure confirmation. Each production
trial includes peer-vote traces from in-character voting by other agents.… See the full description on the dataset page: https://huggingface.co/datasets/ranausmans/synthetic-social-networks.synthetic-pre1930-sftTL;DR
A vintage finetuning dataset (~416k rows, eleven task routes). Sourced by taking excerpts
from pre-1930's texts, turning these into verbatim answers, and then using deepseek-chat to
generate period-appropriate questions of those answers. Any model tuned on this dataset should,
theoretically, never update its weights on anachronistic text, since questions are masked in the
finetuning stages. Features composition, verse, narrative, reasoning, multiturn dialogue, and
calibrated uncertainty… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/synthetic-pre1930-sft.nemotron_synthetic_1T_exp
If you use this project in your research please cite:
@article{patel2025fineinstructions,
title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale},
author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris},
year = {2026},
month = jan,
day = {28},
}
synthetic-mvtec-ad-defect-detection
Synthetic MVTec AD – Defect Detection Dataset by AnywayLabs.ai
Need a custom synthetic dataset for your own defect detection use case?
This dataset is an open-source sample of our synthetic data generation work at AnywayLabs.
If you're working on:
industrial defect detection
visual inspection
supervised anomaly detection
hard-to-collect defect classes
synthetic data for computer vision training
You can request a custom synthetic dataset here, or email:… See the full description on the dataset page: https://huggingface.co/datasets/anywaylabs/synthetic-mvtec-ad-defect-detection.oac-clinical-transport-observability-synthetic
OAC Clinical Transport Observability — Synthetic
This dataset contains 1,200 fixed-seed, entirely synthetic operational
transport-health examples for the companion OAC System Health v1 model.
It contains no records collected from a patient, laboratory, analyzer,
instrument, LIS, EHR, network, or health-care site.
Companion model: OAC System Health v1.
Canonical source: szl-forge clinical gateway.
Data boundary
The closed schema contains only eight bounded… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/oac-clinical-transport-observability-synthetic.clinical-synthetic-text-kg
Data Description
We release the synthetic data generated using the method described in the paper Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models
(ACL 2024 Findings). The external knowledge we use is based on external knowledge graphs.
Generated Datasets
The original train/validation/test data, and the generated synthetic training data are listed as follows. For each dataset, we generate 5000 synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ritaranx/clinical-synthetic-text-kg.unit-price-evidence-synthetic
Unit Price Evidence: Synthetic
This dataset contains rendered synthetic shopping pages and evidence-pointer
targets for product-card discovery and unit-price field extraction. It was built
to warm-start small encoder-decoder models without redistributing retailer HTML,
screenshots, product data, account data, or browsing history.
Release
Version: 0.1.0
Source code: erichasinternet/apples-to-apples
Source manifest SHA-256:… See the full description on the dataset page: https://huggingface.co/datasets/hotdogsalesman/unit-price-evidence-synthetic.mitre-attack-synthetic-scenarios
MITRE ATT&CK Synthetic Scenario Logs v3.0
Expanded Dataset: 30 scenarios × 8 events = 240 synthetic events
Axis
Coverage
Environment
endpoint, cloud, SaaS, identity, CI/CD, OT/IoT
Actor Type
external_apt, ransomware, insider, compromised_vendor, careless_admin, automated_threat
Intent
exfiltration, impact, fraud, persistence, reconnaissance, cryptomining, espionage
Detection Source
EDR, IAM, SIEM, DLP, DNS, proxy, cloud_audit, email_gateway, CASB, NDR, PAM, firewall… See the full description on the dataset page: https://huggingface.co/datasets/koushikcs09/mitre-attack-synthetic-scenarios.mobtranslate-kuku-yalanji-synthetic-corpus-v2
MobTranslate Kuku Yalanji Synthetic Research Corpus v2
Complete public research release of 20,047 synthetic English-Kuku Yalanji sentence pairs plus the process evidence
needed to inspect their production, review, revision, split, and use in the MobTranslate model program.
Project-reviewed synthetic research material pending fluent-speaker and elder verification. It is not a
speaker-certified dictionary or translation corpus.
Identity
ISO 639-3: gvn
Glottocode:… See the full description on the dataset page: https://huggingface.co/datasets/ajaxdavis/mobtranslate-kuku-yalanji-synthetic-corpus-v2.meddeid-english-synthetic-benchmark
MedDeID English synthetic clinical benchmark
This is a fixed, human-validated benchmark containing 300 synthetic English
clinical documents: 150 en-GB and 150 en-US. It contains 1,717 primary PII
spans and 7,358 confirmed core-PII subannotation segments. It contains no real
patient notes or personal information.
Use the entire test split only for final evaluation:
from datasets import load_dataset
benchmark = load_dataset(
"stighellemans/meddeid-english-synthetic-benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/stighellemans/meddeid-english-synthetic-benchmark.Synthetic-Dataset-Childrens-Stories**Status: released 13-09-2026, repacked 14-09-2026.** The 14-09-2026 repack replaced 58
items after the acceptance gates were strengthened (prompt-instruction leaks, markdown
bullet lists and blockquotes); the other 29,942 are unchanged. Development stopped, pipeline released 17/09/26.
SAMPLE RELEASE: 30,000 synthetic children's short stories for early-reader language modelling.
Metrics
Value
genres
26
stories per genre
1.153-1.154K stories
total characters
38… See the full description on the dataset page: https://huggingface.co/datasets/ContextReq/Synthetic-Dataset-Childrens-Stories.meddeid-dutch-synthetic-benchmark
MedDeID Dutch synthetic benchmark
This repository contains the fixed 300-document synthetic Dutch evaluation
benchmark. It contains no real patient notes and must not be mixed into a
training or validation partition when reporting MedDeID benchmark results.
This is the openly shareable synthetic benchmark described in the manuscript. It is not
the separate 300-note hospital benchmark, which contains personal information
and is not publicly distributed.
Subannotations… See the full description on the dataset page: https://huggingface.co/datasets/stighellemans/meddeid-dutch-synthetic-benchmark.OpenSeek-Synthetic-Reasoning-Data-Examples
OpenSeek-Reasoning-Data
OpenSeek [Github|Blog]
Recent reseach has demonstrated that the reasoning ability of LLMs originates from the pre-training stage, activated by RL training. Massive raw corpus containing complex human reasoning process, but lack of generalized and effective synthesis method to extract these reasoning process.
News
🔥🔥🔥[2025/02/25] We publish some math, code, and general knowledge domain reasoning data synthesized from the current pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/OpenSeek-Synthetic-Reasoning-Data-Examples.chatml-synthetic-2026-0420pstu-synthetic-secrets
PSTU Synthetic Secrets Dataset
Synthetic secrets benchmark for evaluating LLM memorization and unlearning, from the paper:
Not All Secrets Are Equal: Type-Aware Unlearning for Language Model Secret Removal
Hoda Fakhar — ECML PKDD 2026
Dataset Description
175 synthetic secrets across 25 types, each paired with 100 structurally similar decoys for computing the Carlini exposure metric.
All data is synthetically generated. No real credentials, PII, or sensitive information… See the full description on the dataset page: https://huggingface.co/datasets/Hodfa71/pstu-synthetic-secrets.syntheticclinical-synthetic-text-llm
Data Description
We release the synthetic data generated using the method described in the paper Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models
(ACL 2024 Findings). The external knowledge we use is based on LLM-generated topics and writing styles.
Generated Datasets
The original train/validation/test data, and the generated synthetic training data are listed as follows. For each dataset, we generate 5000… See the full description on the dataset page: https://huggingface.co/datasets/ritaranx/clinical-synthetic-text-llm.meddeid-english-synthetic-corpus
MedDeID English synthetic clinical corpus
This repository contains 6,700 synthetic English clinical documents with
character-offset de-identification spans: 3,350 en-GB and 3,350 en-US
documents. It contains no real patient notes or personal information. The
complete corpus was used to train
meddeid-english-synth.
Split policy
All records are exposed in one train split. There is no publisher-defined
validation split. Users who tune a model must create and report… See the full description on the dataset page: https://huggingface.co/datasets/stighellemans/meddeid-english-synthetic-corpus.unclickbait-synthetic-27b-trajectoriespie-synthetic
PIE synthetic dataset
Repo: https://github.com/awasthiabhijeet/PIE
Paper: https://aclanthology.org/D19-1435.pdf
slm-synthetic-distillation-sft
SLM Synthetic Distillation SFT
Summary
Synthetic response-distillation dataset containing prompt-response examples generated with a teacher model.
Dataset
Dataset type: response distillation
Total rows: 14,717
Language: English
Signal Distribution
Signal
Rows
arithmetic
2,153
cloud
842
code
994
data_transform
1,078
database
1,190
debugging
1,331
educational_qa
2,124
factual_restraint
2,050
instruction
1… See the full description on the dataset page: https://huggingface.co/datasets/tohio/slm-synthetic-distillation-sft.
