datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthesized_datasetJiuzhang3.0_synthsea-syntheticsynthtraces
SynthTraces
A minimal codebase to generate synthetic coding agent session traces using Pi.
Each session pairs two models working inside one of the project codebases:
a remotely hosted open model (e.g. deepseek-ai/DeepSeek-V4-Pro, openai/gpt-oss-120b, Qwen/Qwen3.6-27B) backs the coding agent, equipped with the default Pi tools — read, write, edit, and bash;
a local model running in llama.cpp plays the user, opening with one of the starting questions and driving the… See the full description on the dataset page: https://huggingface.co/datasets/julien-c/synthtraces.africa-synth-aid-flows-medical-multimodal-fracture-all
Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.unclickbait-synthetic-27b-trajectories
Unclickbait Synthetic 27B Trajectories
Synthetic trajectory dataset generated by using rich structured JSON prompts and validated by two-stage judging pipeline.
Contents
: Full generated trajectories (current snapshot: 48,623 records out of 152,369 pristine event candidates).
: 30 benchmark test samples audited end-to-end through the 122B two-stage judge (Stage 1 integrity gate + Stage 2 4D scoring).
synth-cc-unfilteredmachinelearninglm-scm-synthetic-tabularml
MachineLearningLM Pretraining Corpus
This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.2026-09-15-da-lowstakes-refresh-synth
2026-09-15-da-lowstakes-refresh-synth
field
value
experiment
da-lowstakes-refresh:716 synthetic conversations selected across immutable recipe phases
date_generated
Origin run timestamps: {"phase_ff86738ea4c0df80": "20260915_022437", "phase_fe147df46f721b6f": "20260915_030534"}; publication date=2026-09-15
constitution
constitutions/claude_distilled_09_principles/constitution.md; SHA256=8e273b472d945aa23efa6236886da5e1171bff2193ee31ff73489ca54c4f0edc
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-15-da-lowstakes-refresh-synth.ft-instruction-synthesizer-collection
Instruction Pre-Training: Language Models are Supervised Multitask Learners (EMNLP 2024)
This repo contains the fine-tuning data collection for the context-based instruction synthesizer used in our paper Instruction Pre-Training: Language Models are Supervised Multitask Learners.
We explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response pairs to pre-train language models. The… See the full description on the dataset page: https://huggingface.co/datasets/instruction-pretrain/ft-instruction-synthesizer-collection.synthetic-pii-function-calling
Dataset Summary
A function calling dataset created by filtering the urchade/synthetic-pii-ner-mistral-v1 dataset.
pi-synthetic
Coding agent session traces for aaaaliou/pi-synthetic
This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-synthetic.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-synthetic.2026-09-10-delib-synth
Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7). 658 rows from the full run; the 50 prompts it rejected were re-run in 2 pass(es) with 8 candidates per round under the amended constitution, recovering 42; 8 remain rejected
field
value
experiment
Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-10-delib-synth.Synthetic-AI-ML-Dataset
Synthetic-AI-ML-Dataset
Synthetic Q&A dataset on AI and Machine Learning
Dataset Details
Metric
Value
Topic
AI and Machine Learning
Total Q&A Pairs
14021
Valid Pairs
14021
Provider/Model
ollama/gpt-oss:120b
Generation Cost
Metric
Value
Prompt Tokens
14,941,957
Completion Tokens
17,159,263
Total Tokens
32,101,220
GPU Energy
12.9628 kWh
Sources
This dataset was generated from 474 scholarly papers:
#… See the full description on the dataset page: https://huggingface.co/datasets/Kylan12/Synthetic-AI-ML-Dataset.oac-clinical-transport-observability-synthetic
OAC Clinical Transport Observability — Synthetic
This dataset contains 1,200 fixed-seed, entirely synthetic operational
transport-health examples for the companion OAC System Health v1 model.
It contains no records collected from a patient, laboratory, analyzer,
instrument, LIS, EHR, network, or health-care site.
Companion model: OAC System Health v1.
Canonical source: szl-forge clinical gateway.
Data boundary
The closed schema contains only eight bounded… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/oac-clinical-transport-observability-synthetic.Synthetic-Causal-Reasoning-50k
🏭 Sovereign Synthetic Reasoning Dataset (400k)
"High-Quality Chain-of-Thought Data at Scale."
📊 Overview
This dataset contains 400,000 synthetic reasoning samples spanning 16 enterprise domains (Finance, Pharma, Legal, Cybersecurity, Supply Chain, etc.).
It was generated using the Sovereign Generator, which produced 1.6 million samples and applied a strict quality filter (Top 25%) to retain only the most logically consistent and complex chains.
Average Quality… See the full description on the dataset page: https://huggingface.co/datasets/davidfoss/Synthetic-Causal-Reasoning-50k.synthetic-beir-dataSynthForensicsSynthForensics: Benchmarking and Evaluating People-Centric Synthetic Video Deepfakes
Official Repository for the SynthForensics (SF) Benchmark
Abstract
Modern T2V/I2V generators synthesize people increasingly hard to distinguish from authentic footage, while current evaluation suites lag: legacy benchmarks target manipulation-based forgeries, and recent synthetic-video benchmarks prioritize scale over realistic human depiction. We introduce SynthForensics, a… See the full description on the dataset page: https://huggingface.co/datasets/SynthForensics/SynthForensics.2026-07-29-synthdoc-approved-constitution-sft
Dialogue dataset: Synthetic SFT corpus generated by synthdoc from the approved constitution, to regenerate the difficult-advice training data against a revised specification at a scale comparable to v1, so that the constitution is the intended difference between the two datasets. 1,443 documents / 1,531,369 Qwen3 tokens across five sub-corpora, matching v1's 1.52M.
Required metadata
field
value
experiment
Synthetic SFT corpus generated by synthdoc from… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-synthdoc-approved-constitution-sft.BILGE-Synthetic-Web
BILGE-Synthetic-Web Dataset
BILGE-Synthetic-Web was created following the methodology presented in the Cosmopedia blog/article.
All content was generated using a 27B-parameter model.
Further details on the methodology are available at:
🔗 https://huggingface.co/blog/cosmopedia
nemotron_synthetic_1T_exp
If you use this project in your research please cite:
@article{patel2025fineinstructions,
title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale},
author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris},
year = {2026},
month = jan,
day = {28},
}
synthetic-social-networks
Synthetic Social Networks (Dataset)
Raw experimental outputs from the Synthetic Social Networks study:
59,776 in-character LLM-agent posts from 528 production trials, and
64,562 posts total when the original pipeline-verification runs are
included. The artifact combines an exploratory stage with a separately
frozen, preregistered 448-trial matched-exposure confirmation. Each production
trial includes peer-vote traces from in-character voting by other agents.… See the full description on the dataset page: https://huggingface.co/datasets/ranausmans/synthetic-social-networks.typed-decisions-synth
Typed Decisions Synth
This is the synthetic dataset I made for Hmm, a small open model that answers questions about your data with probabilities instead of text.
It has 7,414 cases with 25,859 questions across 149 domains and workflows. Every question has an answer and a soft label (a probability for every option), so you can train a model to be unsure when it should be.
Code and the model: github.com/n4ze3m/hmm
Note: Everything here is written and labelled by an LLM. Nobody… See the full description on the dataset page: https://huggingface.co/datasets/n4ze3m/typed-decisions-synth.synthetic-mvtec-ad-defect-detection
Synthetic MVTec AD – Defect Detection Dataset by AnywayLabs.ai
Need a custom synthetic dataset for your own defect detection use case?
This dataset is an open-source sample of our synthetic data generation work at AnywayLabs.
If you're working on:
industrial defect detection
visual inspection
supervised anomaly detection
hard-to-collect defect classes
synthetic data for computer vision training
You can request a custom synthetic dataset here, or email:… See the full description on the dataset page: https://huggingface.co/datasets/anywaylabs/synthetic-mvtec-ad-defect-detection.synthetic-pre1930-sftTL;DR
A vintage finetuning dataset (~416k rows, eleven task routes). Sourced by taking excerpts
from pre-1930's texts, turning these into verbatim answers, and then using deepseek-chat to
generate period-appropriate questions of those answers. Any model tuned on this dataset should,
theoretically, never update its weights on anachronistic text, since questions are masked in the
finetuning stages. Features composition, verse, narrative, reasoning, multiturn dialogue, and
calibrated uncertainty… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/synthetic-pre1930-sft.FacturaRD-Synth
Facturas DGII sintéticas
Dataset de facturas dominicanas completamente sintéticas para entrenamiento y
evaluación de extracción fiscal y OCR con modelos multimodales como Florence-2.
El objetivo es entrenar modelos capaces de recibir una imagen con una o varias
facturas y producir simultáneamente:
registros fiscales estructurados;
una transcripción OCR del contenido visible.
Formato
Cada fila de train.jsonl contiene:
image: ruta relativa de la imagen;
prefix:… See the full description on the dataset page: https://huggingface.co/datasets/puruchinera/FacturaRD-Synth.SynthPAI
Dataset Card for SynthPAI
SynthPAI was created to provide a dataset that can be used to investigate the personal attribute inference (PAI) capabilities of LLM on online texts. Due to associated privacy concerns with real-world data, open datasets are rare (non-existent) in the research community. SynthPAI is a synthetic dataset that aims to fill this gap.
Dataset Details
Dataset Description
SynthPAI was created using 300 GPT-4 agents seeded with individual… See the full description on the dataset page: https://huggingface.co/datasets/RobinSta/SynthPAI.clinical-synthetic-text-kg
Data Description
We release the synthetic data generated using the method described in the paper Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models
(ACL 2024 Findings). The external knowledge we use is based on external knowledge graphs.
Generated Datasets
The original train/validation/test data, and the generated synthetic training data are listed as follows. For each dataset, we generate 5000 synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ritaranx/clinical-synthetic-text-kg.unit-price-evidence-synthetic
Unit Price Evidence: Synthetic
This dataset contains rendered synthetic shopping pages and evidence-pointer
targets for product-card discovery and unit-price field extraction. It was built
to warm-start small encoder-decoder models without redistributing retailer HTML,
screenshots, product data, account data, or browsing history.
Release
Version: 0.1.0
Source code: erichasinternet/apples-to-apples
Source manifest SHA-256:… See the full description on the dataset page: https://huggingface.co/datasets/hotdogsalesman/unit-price-evidence-synthetic.mitre-attack-synthetic-scenarios
MITRE ATT&CK Synthetic Scenario Logs v3.0
Expanded Dataset: 30 scenarios × 8 events = 240 synthetic events
Axis
Coverage
Environment
endpoint, cloud, SaaS, identity, CI/CD, OT/IoT
Actor Type
external_apt, ransomware, insider, compromised_vendor, careless_admin, automated_threat
Intent
exfiltration, impact, fraud, persistence, reconnaissance, cryptomining, espionage
Detection Source
EDR, IAM, SIEM, DLP, DNS, proxy, cloud_audit, email_gateway, CASB, NDR, PAM, firewall… See the full description on the dataset page: https://huggingface.co/datasets/koushikcs09/mitre-attack-synthetic-scenarios.
