datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
unclickbait-synthetic-27b-trajectories
Unclickbait Synthetic 27B Trajectories
Synthetic trajectory dataset generated by using rich structured JSON prompts and validated by two-stage judging pipeline.
Contents
: Full generated trajectories (current snapshot: 48,623 records out of 152,369 pristine event candidates).
: 30 benchmark test samples audited end-to-end through the 122B two-stage judge (Stage 1 integrity gate + Stage 2 4D scoring).
machinelearninglm-scm-synthetic-tabularml
MachineLearningLM Pretraining Corpus
This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.synthetic-pii-function-calling
Dataset Summary
A function calling dataset created by filtering the urchade/synthetic-pii-ner-mistral-v1 dataset.
pi-synthetic
Coding agent session traces for aaaaliou/pi-synthetic
This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-synthetic.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-synthetic.synthetic-pre1930-sftTL;DR
A vintage finetuning dataset (~416k rows, eleven task routes). Sourced by taking excerpts
from pre-1930's texts, turning these into verbatim answers, and then using deepseek-chat to
generate period-appropriate questions of those answers. Any model tuned on this dataset should,
theoretically, never update its weights on anachronistic text, since questions are masked in the
finetuning stages. Features composition, verse, narrative, reasoning, multiturn dialogue, and
calibrated uncertainty… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/synthetic-pre1930-sft.Synthetic-Dataset-Childrens-Stories**Status: released 13-09-2026, repacked 14-09-2026.** The 14-09-2026 repack replaced 58
items after the acceptance gates were strengthened (prompt-instruction leaks, markdown
bullet lists and blockquotes); the other 29,942 are unchanged. Development stopped, pipeline released 17/09/26.
SAMPLE RELEASE: 30,000 synthetic children's short stories for early-reader language modelling.
Metrics
Value
genres
26
stories per genre
1.153-1.154K stories
total characters
38… See the full description on the dataset page: https://huggingface.co/datasets/ContextReq/Synthetic-Dataset-Childrens-Stories.pstu-synthetic-secrets
PSTU Synthetic Secrets Dataset
Synthetic secrets benchmark for evaluating LLM memorization and unlearning, from the paper:
Not All Secrets Are Equal: Type-Aware Unlearning for Language Model Secret Removal
Hoda Fakhar — ECML PKDD 2026
Dataset Description
175 synthetic secrets across 25 types, each paired with 100 structurally similar decoys for computing the Carlini exposure metric.
All data is synthetically generated. No real credentials, PII, or sensitive information… See the full description on the dataset page: https://huggingface.co/datasets/Hodfa71/pstu-synthetic-secrets.Turnstile-Synthetic-Domains
Data Turnstile — Synthetic Domains
A large-scale synthetic dataset of function-calling interactions with chain-of-thought reasoning traces, designed for training small language models on tool-use tasks.
Dataset Summary
Metric
Value
Interactions
100,262
Unique APIs
1,025
Distractors per interaction
5
Template types
17
Avg roles per interaction
~10
Avg tokens per interaction
~972
Language
English
Generator model
Qwen2.5-32B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/amazon/Turnstile-Synthetic-Domains.fable-5-coding-and-debugging-traces-synthetic-corrections
Model Synthetic Corrections
1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset.
Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.Persian-Synthetic-Instruct
Persian Synthetic Instruct
High-quality Persian instruction-following dataset generated using LLMs.
4,000+ instruction-response pairs
51 domains
Generated by gpt-4.1-mini and gpt-4.1-nano
math-dow-mod-synthetic-v1
Math + Cyclic-Time Synthetic Dataset
Synthetic dataset for training a small (~10M-100M param), task-specialized
LLM on arithmetic (addition, multiplication), cyclic time arithmetic
(days-of-week, months, 12-hour and 24-hour clock), and mod-k remainder
probes — generalization-focused rather than memorization, following on
from the T2/T5/T10 modular-circuit discussion. Days-of-week, months, and
hours are all instances of the same underlying cyclic/modular-addition
structure… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/math-dow-mod-synthetic-v1.mimo-coding-synthetic-5k
MiMo Coding Synthetic 5.4K
MiMo Coding Synthetic 5.4K is a purely synthetic coding instruction dataset generated with Xiaomi MiMo mimo-v2.5-pro.
It contains 5,411 validated examples across programming languages, coding task types, and difficulty levels. The dataset is provided in two formats:
A canonical rich JSONL format with metadata and labels.
An OpenAI chat messages JSONL format for supervised fine-tuning pipelines.
The generation run used 20 parallel workers for roughly… See the full description on the dataset page: https://huggingface.co/datasets/himanshunakrani9/mimo-coding-synthetic-5k.ALIA-es-legal-administrative-synthetic-instructions
Dataset Introduction
The ALIA Spanish Legal and Administrative Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in legal and administrative tasks with controlled formats and large-scale supervision.
It contains:
763,804 instances
534,112,398 tokens
16 task modalities (questions, instructions, multiple-choice, true/false; with and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-synthetic-instructions.ALIA-es-biomedical-synthetic-instructions
Dataset Introduction
The ALIA Spanish Biomedical Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in biomedical and healthcare tasks with controlled formats and large-scale supervision.
It contains:
639,456 instances
961,073,205 tokens
14 task modalities (clinical diagnosis, patient education, ethical reasoning, document… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-synthetic-instructions.math-dow-mod-synthetic-v1-base6
Math + Cyclic-Time Synthetic Dataset
Synthetic dataset for training a small (~10M-100M param), task-specialized
LLM on arithmetic (addition, multiplication), cyclic time arithmetic
(days-of-week, months, 12-hour and 24-hour clock), and mod-k remainder
probes — generalization-focused rather than memorization, following on
from the T2/T5/T10 modular-circuit discussion. Days-of-week, months, and
hours are all instances of the same underlying cyclic/modular-addition
structure… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/math-dow-mod-synthetic-v1-base6.math-dow-mod-synthetic-v1-base7
Math + Cyclic-Time Synthetic Dataset
Synthetic dataset for training a small (~10M-100M param), task-specialized
LLM on arithmetic (addition, multiplication), cyclic time arithmetic
(days-of-week, months, 12-hour and 24-hour clock), and mod-k remainder
probes — generalization-focused rather than memorization, following on
from the T2/T5/T10 modular-circuit discussion. Days-of-week, months, and
hours are all instances of the same underlying cyclic/modular-addition
structure… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/math-dow-mod-synthetic-v1-base7.synthetic-mental-health-convos
Synthetic Mental Health SFT Dataset
Dataset Summary
This dataset contains high-fidelity, synthetic patient-therapist dialogues designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) in the domain of mental health.
The primary goal of this dataset is to train AI assistants to transition from "general knowledge" models to empathetic, supportive, and safety-conscious mental health companions. The dialogues cover a wide spectrum of mental health conditions… See the full description on the dataset page: https://huggingface.co/datasets/hllzmz/synthetic-mental-health-convos.ALIA-es-cultural-heritage-synthetic-instructions
Dataset Introduction
The ALIA Spanish Cultural and Heritage Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in cultural heritage, digital humanities, and historical knowledge tasks with natural linguistic variation and large-scale supervision.
It contains:
748,480 instances
629,682,398 tokens
25 task modalities (heritage QA… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-synthetic-instructions.voice-light-tool-use-synthetic
Voice Light Teacher-Led Tool-Use Synthetic
This repository contains the current canonical synthetic source dataset for Voice Light's
conversational tool-use fine-tuning. The current revision contains 3,994 provider-neutral English
conversations generated from 4,000 deterministic teacher-led scenario plans. Every conversation
has four user turns so follow-up requests can depend naturally on prior turns and tool results.
The Hugging Face train split names the canonical JSONL file… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/voice-light-tool-use-synthetic.function-calling-synthetic-2000
Synthetic Multi-Turn Function-Calling Conversations
Synthetic, multi-turn function-calling (tool-use) conversations for fine-tuning and evaluating LLMs.
Generated and validated with synthfc
— the open-source pipeline (sampler, prompt builder, validator, post-processor, web viewer) lives in that GitHub repo.
A strong teacher LLM (Qwen/Qwen3.6-35B-A3B) produces each conversation from controlled, sampled
parameters, so the dataset is diverse along many axes (call type, languages… See the full description on the dataset page: https://huggingface.co/datasets/pierjoe/function-calling-synthetic-2000.column-arithmetic-ru-synthetic
Column Arithmetic RU Dataset
Синтетический датасет для обучения модели сложению и вычитанию в столбик.
Splits
train.jsonl: основное обучение
eval.jsonl: holdout-оценка
hard.jsonl: трудные случаи с длинными переносами и займами
Hard cases included
9999+1
10000+9999
9090+1010
55555+55555
10999+2
1234+8766
1000-7
10000-9999
50005-49999
8000-1
10101-909
100000-1
99009+991
12000-3456
700000+300001
1002003-998877
Current release status… See the full description on the dataset page: https://huggingface.co/datasets/foxycuter/column-arithmetic-ru-synthetic.synthetic-math
Synthetic MATH Dataset
Dataset Summary
This dataset contains symthetic math problems generated with GPT-4o and verified with DeepSeek-R1, intended to augment the MATH dataset (Hendrycks et al., 2021) with additional training/evaluation examples.
Only problems where R1's final answer matched the reference answer given by GPT-4o are included. Each row bundles the problem, the reference solution, and R1's full reasoning trajectory used for verification.
Note: This is… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/synthetic-math.turkish-synthetic-corpus
Turkish Synthetic Corpus
A synthetic Turkish text corpus with 1,871,131 documents, designed for Turkish language model training.
About
Inspired by HuggingFaceTB/smollm-corpus. Questions and prompts were sourced from the SmolLM Corpus pipeline; a language model then generated localized Turkish responses and documents around them. All credit for the original corpus design and methodology goes to the HuggingFace SmolLM team.
The resulting dataset covers a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/cturan/turkish-synthetic-corpus.snh-loan-adjudication-synthetic
SNH Synthetic Loan Adjudication Dialogues
This dataset was created for a technical coding challenge about conversational personal-loan adjudication. Every record contains supplied policy rules, a synthetic dialogue, deterministic ground-truth fields, a decision, failed rule IDs, and a customer-facing explanation.
No record represents a real person, and the dataset contains no real applicant information.
Splits
File
Records
Purpose
train.jsonl
4,000… See the full description on the dataset page: https://huggingface.co/datasets/sabber/snh-loan-adjudication-synthetic.Apple-Synthetic
Dataset Summary
A synthetic question–answer dataset grounded in a curated set of seed documents that reflect Apple's historical design philosophy and cultural principles. Questions are interpretive and scenario-based; answers are required to be derivable from the provided source text [Apple-legacy-corpus](https://huggingface.co/datasets/SP4ND4N/Apple-legacy-corpus.
Source type: synthetic, generated from internal seed docs (not scraped Apple manuals)
Format: JSON Lines (JSONL) with… See the full description on the dataset page: https://huggingface.co/datasets/SP4ND4N/Apple-Synthetic.Synthetic-Japanese-Roleplay-NSFW-Claude-3.5s-15.3k-formatted
Synthetic-Japanese-Roleplay-NSFW-Claude-3.5s-15.3k-formatted
20240907 データ増量(約10500件→約15300件)
概要
Claude 3.5 Sonnetを用いて作成した日本語ロールプレイデータセットであるAratako/Synthetic-Japanese-Roleplay-NSFW-Claude-3.5s-15.3kにsystem messageを追加して整形したデータセットです。
データの詳細については元データセットのREADMEを参照してください。
ライセンス
CC-BY-NC-SA 4.0の元配布します。
また、Anthropicの利用規約に記載のある通り、このデータを使ってAnthropicのサービスやモデルと競合するようなモデルを開発することは禁止されています。
seed-free-synthetic-instruct-thai-v1
Seed-Free Synthetic Instruct Thai v1 (F+C+D+)
This dataset is part of the research paper "Seed-Free Synthetic Data Generation Framework for Instruction-Tuning LLMs: A Case Study in Thai" submitted to ACL SRW 2024. It represents the best-performing synthetic dataset (F+C+D+) generated using our novel seed-free framework for low-resource languages, specifically Thai.
Dataset Details
Size: 5,000 instructions
Language: Thai
Task: Instruction-tuning for Large Language Models… See the full description on the dataset page: https://huggingface.co/datasets/parinzee/seed-free-synthetic-instruct-thai-v1.miobra-synthetic-construction-material-extraction-v1
Mi Obra Synthetic Construction Material Extraction v1.0
English
Dataset Description
This dataset contains 10,000 synthetic Spanish construction-material titles
paired with structured entity-extraction targets. It is designed for Argentine
construction terminology and controlled experiments in fine-tuning, teaching,
information extraction, and structured generation.
Language: Spanish (es)
Regional context: Argentina
Rows: 10,000
License: CC BY 4.0… See the full description on the dataset page: https://huggingface.co/datasets/mi-obra/miobra-synthetic-construction-material-extraction-v1.Synthetic-archive
The Synthetic Archive
Synthetic Archive is a large synthetic English-text dataset generated from OCR-derived historical and period-style passages. Knowledge cutoff is year 1900.
Each source passage was divided into chunks and processed through several generation tasks, including:
generating continuations of unfinished passages;
creating question-and-answer pairs;
extracting and reformulating factual knowledge;
rewriting material as a narrative;
transforming source material into… See the full description on the dataset page: https://huggingface.co/datasets/croqaz/Synthetic-archive.agentic_synthetic_aggressive_conversations_en
Simulated Aggressive Customer Service Conversations Dataset
Overview
This dataset contains aggressive customer service conversations generated by an agentic simulation system.
Each record is stored in JSON Lines (JSONL) format and includes:
Scenario Metadata: Selected bank, customer, agent profiles, and task details.
Conversation Messages: Full message history between the customer and service agent.
Summary: A German summary of the conversation.
Cost Metrics: API cost… See the full description on the dataset page: https://huggingface.co/datasets/marccgrau/agentic_synthetic_aggressive_conversations_en.
