datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ID_Legal_QA_SynThink
🧠 Indonesian Legal QA Synthetic Think Dataset (ID_Legal_QA_SynThink)
This repository features an advanced Synthetic Question-and-Answer dataset for the Indonesian legal domain, distinguished by the inclusion of an explicit Thinking Phase (Chain-of-Thought). 🏛️
💡 The Concept: Transparent Legal Reasoning
Standard QA datasets often provide just the "final answer." This dataset goes deeper by capturing the internal reasoning process of the model before it arrives at a… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_Legal_QA_SynThink.synthea-575k-patients
Synthea Synthetic Patient Records (575K Patients)
A comprehensive synthetic healthcare dataset containing 575,415 patients with complete medical histories, generated using Synthea — the gold standard for synthetic EHR data.
No real patient data. Fully synthetic, HIPAA-safe, and ready for ML research and education.
Why This Dataset?
575K patients with realistic demographics, conditions, medications, and encounters
Privacy-safe: No real PHI — use freely in research… See the full description on the dataset page: https://huggingface.co/datasets/richardyoung/synthea-575k-patients.recursive-task-synthesis-glm-5.3-rollouts
GLM 5.3 agentic rollouts on Recursive-Task-Synthesis
This dataset catalogs the full collection made from the pinned
Recursive-Task-Synthesis dataset revision
be44f96808d5a9b599d5cb024341ff00091adeb7. The repository includes approximately 260.5 GiB of trajectory payload tar shards.
Contents at a glance
Item
Count
Source tasks considered
37,284
Source candidates inspected
19,368
Converted tasks after source filters
18,600
Tasks passing gold… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/recursive-task-synthesis-glm-5.3-rollouts.pi-synthetic
Coding agent session traces for aaaaliou/pi-synthetic
This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-synthetic.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-synthetic.synthesis
NuBerea/synthesis
A cross-corpus synthesis layer for the study of early Jewish and Christian literature.
Each config joins pericope-level text units from one corpus — the canonical Bible
(Old and New Testament), Second Temple Pseudepigrapha, the Aramaic Targumim, the Nag
Hammadi corpus, or Greek and Latin patristic authors — with rhetorical claims extracted
from those units and with links into a shared concept vocabulary. The result is a set
of per-corpus tables that let a… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/synthesis.synthetic-pre1930-sftTL;DR
A vintage finetuning dataset (~416k rows, eleven task routes). Sourced by taking excerpts
from pre-1930's texts, turning these into verbatim answers, and then using deepseek-chat to
generate period-appropriate questions of those answers. Any model tuned on this dataset should,
theoretically, never update its weights on anachronistic text, since questions are masked in the
finetuning stages. Features composition, verse, narrative, reasoning, multiturn dialogue, and
calibrated uncertainty… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/synthetic-pre1930-sft.hiring-bias-mitigation-synthetic-data
Hiring-bias mitigation — synthetic training data
Semi-synthetic data for training LLMs to make hiring decisions that do not depend on a
protected attribute (military status, gender, religion), in English and Ukrainian.
Real inputs, synthetic labels. CVs and job descriptions are real, anonymised postings
from the Djinni Recruitment Dataset (MIT). Decisions and rationales were written by the
teacher model Qwen/Qwen3.5-122B-A10B-GPTQ-Int4.
Code and results:… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/hiring-bias-mitigation-synthetic-data.MedFact-SynthTo train Med-V1, we construct MedFact-Synth, a large-scale synthetic training set including 1.5 million instances.
Each instance contains: a synthetic claim to be verified, a source article serving as evidence, a rationale explaining the verification, and a 5-point Likert-scale verdict, ranging from strong contradiction (-2) and partial contradiction (-1) to neutral (0), partial agreement (+1), and strong agreement (+2).
To build this dataset, we begin by sampling one million articles from… See the full description on the dataset page: https://huggingface.co/datasets/nlm-dir/MedFact-Synth.SyntheticTalk-jp
LambdaTalk-v2
LambdaTalk-v2 is a Japanese synthetic multi-turn conversation dataset generated with Gemma 4 31B. It contains conversations based on seed questions collected from 36 source datasets.
Each conversation contains three user-assistant turns. The first user message is the original seed question. The remaining five messages were generated by Gemma 4 31B.
Purpose
The main purpose of this dataset is supervised fine-tuning of Japanese conversational language… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/SyntheticTalk-jp.tau-bench-synthetic
tau-bench-synthetic
Synthetic tool-use training data for tau-bench, generated using a GT-first task construction pipeline with GLM-5 (via Fireworks API) as the trajectory generator.
Overview
This dataset was built to train small LLMs (e.g., Qwen3-1.7B) on multi-turn tool-use tasks without using the original tau-bench evaluation set. The pipeline follows a GT-first approach: ground-truth actions are constructed programmatically from the database, then an LLM generates… See the full description on the dataset page: https://huggingface.co/datasets/fuvty/tau-bench-synthetic.synthetic-persian-v1
IbnSina Synthetic Persian Corpus v1
Sina Meraji · ORCID 0009-0002-8028-1932 · github.com/sinameraji
پیکرهٔ مصنوعی فارسی ابنسینا (نسخهٔ ۱) — ۲٫۰۷۵ میلیارد توکن متن آموزشیِ فارسی که از ابتدا به فارسی تولید شده است، نه ترجمه از انگلیسی. این پیکره برای پوشش حوزههایی ساخته شده که وبِ فارسی در آنها کممایه است: توضیح مفاهیم علمی و مهندسی، مسئلههای حلشدهٔ ریاضی و فیزیک، زنجیرههای استدلال، و متنهای علمی-پزشکی. هر سند را یک داور خودکار با معیارهای سختگیرانه (درستیِ محاسبهها،… See the full description on the dataset page: https://huggingface.co/datasets/ibnsina-llm/synthetic-persian-v1.Synthetic-JP-EN-Coding-Dataset-801k
Synthetic-JP-EN-Coding-Dataset-801k
Magpieによって作成したコードSFTデータセットであるAratako/Synthetic-JP-EN-Coding-Dataset-Magpie-69kを元に、Evol-Instructのような手法を用いて複数のinstructionとresonseを生成し拡張して作成した、日英混合801262件のコードSFT用合成データセットです。
日本語: 173849件
英語: 627413件
元のinstructionの作成に利用したモデルは以下の通りです。modelキーに該当レコードの作成に利用したモデル情報があります。
nvidia/Nemotron-4-340B-Instruct
microsoft/Phi-3-medium-4k-instruct
mistralai/Mixtral-8x22B-Instruct-v0.1… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-JP-EN-Coding-Dataset-801k.SyntheticTextbook-jp
SyntheticTextbook-jp
SyntheticTextbook-jp is a Japanese synthetic text dataset generated with Gemma 4 26B and Gemma 4 31B. The dataset was created by rewriting noisy source text into textbook-style Japanese for elementary school, junior high school, and high school levels. The rewritten text keeps only general knowledge from the source text.
Purpose
The main purpose of this dataset is to help LLMs learn natural Japanese text flow. This dataset is designed around… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/SyntheticTextbook-jp.fable-5-coding-and-debugging-traces-synthetic-corrections
Model Synthetic Corrections
1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset.
Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.synthweb-gemma4-26b-a4b
Gemma-4-26B-A4B FineWeb Rollouts (~580k docs)
Open-ended continuations of FineWeb
(sample-10BT) document prefixes, generated by google/gemma-4-26b-a4b (the base, non-it
Gemma-4 26B-A4B mixture-of-experts model), then mode-collapse filtered. This is the Gemma-4
analogue of cds-jb/qwen3-8b-fineweb-rollouts-100k:
a "synthweb" corpus of natural model-generated documents, intended as the substrate for
activation-oracle / interpretability probing (extract a base model's residual… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/synthweb-gemma4-26b-a4b.GNOTHEIA-synthetic-insurance-dataset
GNOTHEIA Synthetic Insurance Dataset
Published by: Gratex International a.s.Project: InnovAIte — InnovAIte Slovakia
License: Apache 2.0Version: 1.0.0Contact: info@gratex.com
A synthetic insurance claims dataset designed for AI systems that evaluate insurance claims using OMG SBVR business rules, structured claim polycontexts and synthetic claim-related documents.
The dataset main goal is to support:
LLM fine-tuning pipeline
SBVR reasoning benchmarks
insurance claim AI… See the full description on the dataset page: https://huggingface.co/datasets/gratex/GNOTHEIA-synthetic-insurance-dataset.synthetic_multilingual_llm_prompts
Image generated by DALL-E. See prompt for more details
📝🌐 Synthetic Multilingual LLM Prompts
Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.tulu-3-sft-personas-if-synth
Tulu3 Personas Synth Reasoning
Synthetic reasoning traces for allenai/tulu-3-sft-personas-instruction-following. Each record contains a persona-based instruction with SYNTH-style reasoning and a generated answer.
Dataset Summary
24,747 records (18 dupes + 4,281 incomplete/truncated removed from 29,046 source)
24,747 reasoning turns (99.9% format compliance)
Average 1,398 chars per reasoning trace
Incomplete records (cut off mid-sentence due to early 1024… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/tulu-3-sft-personas-if-synth.synthid-qwen3-4b-instruct-2507-wildchat
Qwen3-4B SynthID three-arm corpus
This export contains aligned unwatermarked, SynthID key-A, and SynthID key-B
responses from Qwen/Qwen3-4B-Instruct-2507. Matched splits share prompts
and request seeds across configurations; unmatched splits use mutually disjoint
prompt pools.
Export complete for its source work queue: true.
Generation profile
Model revision: cdbee75f17c01a7cc42f958dc650907174af0554
Native model dtype: bfloat16
Maximum generated tokens: 4096… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/synthid-qwen3-4b-instruct-2507-wildchat.synthetic-misconceptions-conversations
Synthetic Misconceptions Conversations
All data in this dataset is synthetic. No conversation here was had by a real
person. The only human-authored source material is Wikipedia text: the corrections
in List of common misconceptions about science, technology, and
mathematics
(260 entries), plus entries from List of conspiracy
theories and
Category:Health-related conspiracy
theories
(85 entries, filtered — see below). All of it is CC BY-SA licensed on
Wikipedia. Everything… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/synthetic-misconceptions-conversations.column-arithmetic-ru-synthetic
Column Arithmetic RU Dataset
Синтетический датасет для обучения модели сложению и вычитанию в столбик.
Splits
train.jsonl: основное обучение
eval.jsonl: holdout-оценка
hard.jsonl: трудные случаи с длинными переносами и займами
Hard cases included
9999+1
10000+9999
9090+1010
55555+55555
10999+2
1234+8766
1000-7
10000-9999
50005-49999
8000-1
10101-909
100000-1
99009+991
12000-3456
700000+300001
1002003-998877
Current release status… See the full description on the dataset page: https://huggingface.co/datasets/foxycuter/column-arithmetic-ru-synthetic.african-history-sft-synthetic
African History SFT (Chat)
A deduplicated, chat-format supervised fine-tuning (SFT) dataset about African
history and culture, assembled from five source datasets and prepared as a
ready-to-train train/test split.
Each row is a multi-turn conversation in the standard messages format
(system / user / assistant), making it directly usable with
tokenizer.apply_chat_template and TRL's SFTTrainer.
Dataset at a glance
Split
Rows
train
25,552
test
1,345… See the full description on the dataset page: https://huggingface.co/datasets/MaatAI/african-history-sft-synthetic.fable-5-coding-and-debugging-traces-synthetic
Model Synthetic Corrections
1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset.
Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/11-47/fable-5-coding-and-debugging-traces-synthetic.synthcot
synthcot — carrier-paired chains-of-thought for CODI latent decoding
Decoding CODI latent thoughts into chains-of-thought while controlling for the prompt-inversion
pathway (a decoder that cheats by reconstructing the problem text from activations and re-solving).
Analog of cds-jb/synthcognition2 for math CoT: every card pairs one computational skeleton
(the "cognition") with 5 GSM8k-style carriers — word problems that all force exactly that
computation — and the decode label is… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/synthcot.database-query-logs-synthetic
Database Query Logs (synthetic)
3,995 database query-log entries spanning 10 engines - MySQL, PostgreSQL, MongoDB, SQL
Server, Oracle, MariaDB, SQLite, Cassandra, Redis, and Elasticsearch - with query text,
type, complexity, execution timing, and row-count metadata.
These queries are synthetic
The queries were programmatically generated, not captured from production systems.
They were produced by templating a set of query shapes across industry-flavored schema… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/database-query-logs-synthetic.synthesized-coding-assistant-dataset
Synthesized Coding Assistant Dataset
Overview
Coding assistants are increasingly used for real-world software engineering workflows. However, there are relatively few datasets that closely resemble how such assistants operate in practice.
Many existing coding datasets are based on single-turn or single-iteration tasks, where a model receives one coding request and directly produces an answer or patch. In contrast, practical coding assistants often work through… See the full description on the dataset page: https://huggingface.co/datasets/squeezebits/synthesized-coding-assistant-dataset.synthetic-abandoned-cart-email-examples
Synthetic Abandoned Cart Email Examples
An entirely synthetic, bilingual collection of abandoned-cart email drafts with transparent checklist annotations. It is intended for education, prototyping, and evaluation, and contains no real recipients, customer messages, orders, merchant data, or campaign results.
Dataset Description
The dataset mirrors the five visible checks in NeuroCheckout's public Abandoned Cart Email Checker:
message clarity;
primary call to… See the full description on the dataset page: https://huggingface.co/datasets/neurocheckout-ai/synthetic-abandoned-cart-email-examples.RevUtil_synthetic
RevUtil: Measuring the Utility of Peer Reviews for Authors
📄 Paper
💻 GitHub Repository
📚 Overview
Providing constructive feedback to authors is a key goal of peer review. To support research on evaluating and generating useful peer review comments, we introduce RevUtil, a dataset for measuring the utility of peer review feedback.
RevUtil focuses on four main aspects of review comments:
Actionability – Can the author act on the comment?
Grounding & Specificity –… See the full description on the dataset page: https://huggingface.co/datasets/boda/RevUtil_synthetic.Strudel-Synth
Strudel-Synth
Strudel-Synth is a synthetic corpus of 21,174 (MIDI, Strudel) pairs for
training and evaluating MIDI-to-Strudel decompilation, introduced in
Decomposer: Learning to Decompile Symbolic Music to Programs.
📄 Paper: arXiv:2607.01849
🌐 Project page: yewon-kim.com/decomposer
🎹 Live demo: haiyewon/decomposer-demo
🤗 Model: haiyewon/Decomposer-Qwen3-8B
💻 Code: github.com/elianakim/Decomposer
Each pair consists of a Strudel program distilled from Claude-Opus-4.6… See the full description on the dataset page: https://huggingface.co/datasets/haiyewon/Strudel-Synth.Heterogenous_Synthesis_Benchmark
Heterogenous_Synthesis_Benchmark
This repository presents a diverse tabular data generation benchmark. We invite you to refer to our paper on arxiv to explore the mechanism behind our data diversity, which we called Distribution-Guided-Rule (DGR). Within this benchmark, you can experience how diverse preference data coverage combined with customized generation enhances post-training performance.
Additionally, Heterogenous_Synthesis_Benchmark includes a comprehensive toolkit for… See the full description on the dataset page: https://huggingface.co/datasets/CurryOvO/Heterogenous_Synthesis_Benchmark.
