datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SYNTH
SYNTH
Blog announcement
SYNTH is the first open generalist synthetic dataset for training small reasoning model end-to-end, jointly released by Pleias and the AI Alliance.
SYNTH includes 79,648,272 individual text samples, comprising over 41 billion words (about 75 billion tokens with Pleias tokenizer). It is based on the amplification of 58,698 articles from Wikipedia and made possible thanks to the Structured Wikipedia dataset from Wikimedia Enterprise.
SYNTH differs… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/SYNTH.Big-Math-RL-Verified
Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
Big-Math is the largest open-source dataset of high-quality mathematical problems, curated specifically for reinforcement learning (RL) training in language models. With over 250,000 rigorously filtered and verified problems, Big-Math bridges the gap between quality and quantity, establishing a robust foundation for advancing reasoning in LLMs.
Request Early Access to Private… See the full description on the dataset page: https://huggingface.co/datasets/SynthLabsAI/Big-Math-RL-Verified.SYNTH
SYNTH
SYNTH is the first open generalist synthetic dataset for training small reasoning model end-to-end, jointly released by Pleias and the AI Alliance.
SYNTH includes 79,648,272 individual text samples, comprising over 41 billion words (about 75 billion tokens with Pleias tokenizer). It is based on the amplification of 58,698 articles from Wikipedia and made possible thanks to the Structured Wikipedia dataset from Wikimedia Enterprise.
SYNTH differs from existing open synthetic… See the full description on the dataset page: https://huggingface.co/datasets/SYNTH-Initiative/SYNTH.context_qa_sum_qwen3_synthetic
Context-based QA and Summarization Synthetic Dataset
Overview
This dataset contains synthetic context-based question-answering (QA) and summarization data. The data was synthesized using:
Source context: openbmb/Ultra-FineWeb
Synthesis model: Qwen3-30B-A3B-Instruct-2507
Each context is obtained by taking the initial segment of raw pretraining text from Ultra-FineWeb, truncated to at most the corresponding number of tokens, while ensuring the truncation does not occur in… See the full description on the dataset page: https://huggingface.co/datasets/yuyijiong/context_qa_sum_qwen3_synthetic.A3-Synth
A3-Synth
💾 Code
📄 Paper
🌐 Website
🤗 Dataset
🤖 Models
📦 PyPI
Structured Distillation of Web Agent Capabilities Enables Generalization
Xing Han Lù, Siva Reddy
A3-Synth is a synthetic training dataset for web agents, generated using the Agent-as-Annotators (A3) framework. It contains ~16k SFT training examples produced by Gemini 3 Pro acting as the Annotator across 3,000 tasks on 6 WebArena environments.
Dataset Structure
A3-Synth/
training/… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/A3-Synth.synthetic_text_to_sql
Image generated by DALL-E. See prompt for more details
synthetic_text_to_sql
gretelai/synthetic_text_to_sql is a rich dataset of high quality synthetic Text-to-SQL samples,
designed and generated using Gretel Navigator, and released under Apache 2.0.
Please see our release blogpost for more details.
The dataset includes:
105,851 records partitioned into 100,000 train and 5,851 test records
~23M total tokens, including ~12M SQL tokens
Coverage across 100 distinct… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_text_to_sql.Synth-APIGen-v0.1
Dataset card for Synth-APIGen-v0.1
This dataset has been created with distilabel.
Pipeline script: pipeline_apigen_train.py.
Dataset creation
It has been created with distilabel==1.4.0 version.
This dataset is an implementation of APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets in distilabel,
generated from synthetic functions. The process can be summarized as follows:
Generate (or in this case modify) python… See the full description on the dataset page: https://huggingface.co/datasets/argilla/Synth-APIGen-v0.1.ID_Legal_QA_SynThink
🧠 Indonesian Legal QA Synthetic Think Dataset (ID_Legal_QA_SynThink)
This repository features an advanced Synthetic Question-and-Answer dataset for the Indonesian legal domain, distinguished by the inclusion of an explicit Thinking Phase (Chain-of-Thought). 🏛️
💡 The Concept: Transparent Legal Reasoning
Standard QA datasets often provide just the "final answer." This dataset goes deeper by capturing the internal reasoning process of the model before it arrives at a… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_Legal_QA_SynThink.unclickbait-synthetic-27b-trajectories
Unclickbait Synthetic 27B Trajectories
Synthetic trajectory dataset generated by using rich structured JSON prompts and validated by two-stage judging pipeline.
Contents
: Full generated trajectories (current snapshot: 48,623 records out of 152,369 pristine event candidates).
: 30 benchmark test samples audited end-to-end through the 122B two-stage judge (Stage 1 integrity gate + Stage 2 4D scoring).
TranslatePsy-AfriSLM-Synthetic-Mix
TranslatePsy-AfriSLM Synthetic Mix
TranslatePsy-AfriSLM Synthetic Mix is a quality-filtered synthetic parallel corpus for machine translation between English and 19 Sub-Saharan African languages. It contains 215,653,192 bidirectional training examples and was selected as the primary African translation component used to post-train the TranslatePsy-AfriSLM model family.
The dataset accompanies the EMNLP 2026 paper TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource… See the full description on the dataset page: https://huggingface.co/datasets/qvac/TranslatePsy-AfriSLM-Synthetic-Mix.latex-data-pub
Hyperion: Scientific LaTeX Corpus (Public)
This dataset constitutes the Public Scientific Corpus for the Hyperion Project at the Synthetix Institute. It contains high-fidelity LaTeX source text extracted from diverse scientific repositories, optimized for topological knowledge discovery and relational mapping.
Dataset Details
Total Documents: ~1,318,468
Average Document Length: Variable (approx. 32KB - 256KB)
Primary Domain: Mathematics, Physics, and Chemistry.
Goal:… See the full description on the dataset page: https://huggingface.co/datasets/synthetix-institute/latex-data-pub.entropy-shard-002
Synthetic Entropy Shards (v2)
Dataset Description
This dataset consists of high-entropy binary shards generated for stress-testing data loading pipelines and training robustness against random noise injection in large-scale tensor operations.
Usage
These files are intended to be consumed as raw byte streams. Due to the stochastic nature of the generation process, the data mimics encrypted traffic patterns or high-density compression artifacts.
Warning: Do not… See the full description on the dataset page: https://huggingface.co/datasets/Synthetic-Entropy-Labs/entropy-shard-002.OpenHand-Synth
Dataset Card for OpenHand-Synth
📜 Paper: OpenHand-Synth: A Large-Scale Synthetic Handwriting Dataset for Multimodal Language Models
Sample Images
Image
Ground Truth
Source
Language
CER
JW
02-10-1436
faker-date
por
0.10
0.96
Stephan Thomsen-Johansen
faker-name
dan
0.0
1.0
Le chat mange.
tatoeba
fra
0.0
1.0
Classical musicsoothes me.She took the risk, knowing that shemight lose a lot of money.I could not catcha single word of their talk.In the old days… See the full description on the dataset page: https://huggingface.co/datasets/to-be/OpenHand-Synth.zenyx-v3-synthetic-sft-v4
Zenyx-V3-Synthetic-SFT-V4 (Master Expansion)
This is the finalized V4 Master Collection for the Zenyx project, expanding upon the previous V3. This version focuses on high-reasoning, code, and math capabilities through massive distillation and thinking-chain integration.
Dataset Summary
Total Samples: 101,523
Branding: Rebranded to Zenyx / Zenyx Lab.
Filtering: Strict English, Math, and Code filter applied. Chinese and non-standard characters removed.
Format:… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/zenyx-v3-synthetic-sft-v4.entropy-shard-001
Synthetic Entropy Shards (v1)
Dataset Description
This dataset consists of high-entropy binary shards generated for stress-testing data loading pipelines and training robustness against random noise injection in large-scale tensor operations.
Usage
These files are intended to be consumed as raw byte streams. Due to the stochastic nature of the generation process, the data mimics encrypted traffic patterns or high-density compression artifacts.
Warning: Do not… See the full description on the dataset page: https://huggingface.co/datasets/Synthetic-Entropy-Labs/entropy-shard-001.Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.lmsys-chat-1m-synth
LMSYS-Chat-1M-Synth: Japanese/English Synthetic Conversation Dataset Derived from LMSYS-Chat-1M
This repository contains a series of Japanese and English conversation datasets derived from LMSYS-Chat-1M.
Llama-3.1-LMSYS-Chat-1M-Synth
Utilized in the post-training of Llama-3.1-Swallow-8B-Instruct-v0.1 and Llama-3.1-Swallow-70B-Instruct-v0.1
Gemma-2-LMSYS-Chat-1M-Synth
Utilized in the post-training of Llama-3.1-Swallow-8B-Instruct-v0.3 and Llama-3.1-Swallow-70B-Instruct-v0.3… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/lmsys-chat-1m-synth.machinelearninglm-scm-synthetic-tabularml
MachineLearningLM Pretraining Corpus
This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.SYNTH-Swallow-Math-Code-Mix
Mixed dataset: SYNTH + SwallowMath-v2 + SwallowCode-v2
This high-signal, all synthetic dataset is a complete shuffled mix of the following four sources:
SYNTH ~63.5%
SwallowCode-v2 ~15.5%
SwallowMath-v2-textbook ~10.5%
SwallowMath-v2-qa ~10.0%
The motivation to provide this on HF was the need for a convenient, pre-shuffled merge of the highest quality synthetic / augmented datasets for small language model pre-training experiments as of… See the full description on the dataset page: https://huggingface.co/datasets/TMoC/SYNTH-Swallow-Math-Code-Mix.latex-data
Hyperion: Scientific LaTeX Corpus
This dataset constitutes the Internal Scientific Corpus for the Hyperion Project of Synthetix Institute. It contains LaTeX source documents used for high-fidelity relational extraction and internal benchmarking of the Epistemic Manifold.
Dataset Details
Total Documents: ~1,192,727 (Internal Base)
Status: Private
Access: Restricted to Synthetix Institute authorized personnel.
Primary Use: Training the private sector of the Epistemic… See the full description on the dataset page: https://huggingface.co/datasets/synthetix-institute/latex-data.openhands-synthetic-conversations
OpenHands Synthetic Conversations
Overview
28 synthetic OpenHands V1 agent conversations generated by running diverse coding-task prompts against the OpenHands SDK with 4 models rotated round-robin. Each conversation captures a complete agentic session: system prompt, user message, tool calls, terminal observations, and the agent's final reply — exactly as produced by the app.all-hands.dev "Download Conversation" export.
Intended use: raw material for indexing /… See the full description on the dataset page: https://huggingface.co/datasets/rajistics/openhands-synthetic-conversations.synthea-575k-patients
Synthea Synthetic Patient Records (575K Patients)
A comprehensive synthetic healthcare dataset containing 575,415 patients with complete medical histories, generated using Synthea — the gold standard for synthetic EHR data.
No real patient data. Fully synthetic, HIPAA-safe, and ready for ML research and education.
Why This Dataset?
575K patients with realistic demographics, conditions, medications, and encounters
Privacy-safe: No real PHI — use freely in research… See the full description on the dataset page: https://huggingface.co/datasets/richardyoung/synthea-575k-patients.synth-apigen-qwen
Dataset Card for argilla-warehouse/synth-apigen-qwen
This dataset has been created with distilabel.
The pipeline script was uploaded to easily reproduce the dataset:
synth_apigen.py.
Dataset creation
This dataset is a replica in distilabel of the framework
defined in: APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets.
Using the seed dataset of synthetic python functions in argilla-warehouse/python-seed-tools,
the… See the full description on the dataset page: https://huggingface.co/datasets/argilla-warehouse/synth-apigen-qwen.synthetic-pii-function-calling
Dataset Summary
A function calling dataset created by filtering the urchade/synthetic-pii-ner-mistral-v1 dataset.
Infinity-Instruct-RU-Synthetic
Infinity-Instruct-RU-Synthetic
A large-scale Russian-language instructional dataset based on Infinity-Instruct by BAAI.
This is not a translation of English answers — it is an independent Russian-language dataset, where only the instructions are sourced from the original set, and all answers are newly generated in Russian from scratch.
To translate the instructions, YandexGPT-5-Lite-8B-instruct was used with a specially fine-tuned LoRA adapter designed for this dataset. The original… See the full description on the dataset page: https://huggingface.co/datasets/KirillR/Infinity-Instruct-RU-Synthetic.recursive-task-synthesis-glm-5.3-rollouts
GLM 5.3 agentic rollouts on Recursive-Task-Synthesis
This dataset catalogs the full collection made from the pinned
Recursive-Task-Synthesis dataset revision
be44f96808d5a9b599d5cb024341ff00091adeb7. The repository includes approximately 260.5 GiB of trajectory payload tar shards.
Contents at a glance
Item
Count
Source tasks considered
37,284
Source candidates inspected
19,368
Converted tasks after source filters
18,600
Tasks passing gold… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/recursive-task-synthesis-glm-5.3-rollouts.UltraData-Math-L3-Textbook-Exercise-Synthetic-split
UltraData-Math L3 Textbook Exercise Synthetic Split
Source dataset: openbmb/UltraData-Math
Source config: UltraData-Math-L3-Textbook-Exercise-Synthetic
Each row contains:
uid
question
answer
The original content field was split using the literal markers
The exercise: and The solution:.
pi-synthetic
Coding agent session traces for aaaaliou/pi-synthetic
This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-synthetic.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-synthetic.staatsblad-synth-nl
Synthetic Dutch from the Belgisch Staatsblad
Diverse, fluent Dutch pretraining text synthesized from
guust-franssens/belgisch-staatsblad
(CC0, Belgian official-gazette filings). Adds Belgium/Flanders coverage to Dutch LM pretraining
mixes, where clean Belgian-Dutch prose is otherwise scarce.
The source text is noisy OCR from scanned PDFs, but its metadata (company, juridical form,
act type, city, date) is clean. A local LLM (google/gemma-2-9b-it)
"launders" the OCR + metadata… See the full description on the dataset page: https://huggingface.co/datasets/pdelobelle/staatsblad-synth-nl.gretel-synthetic-text-to-sql
Fork of gretelai/synthetic_text_to_sql
The gretelai/synthetic_text_to_sql dataset is a large, Apache 2.0 licensed, synthetic Text-to-SQL dataset consisting of 105,851 high-quality records across 100 diverse domains, designed for training language models. It includes comprehensive SQL tasks with varying complexities, database contexts, natural language explanations, and contextual tags, outperforming existing datasets in SQL correctness and standards compliance.
