datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-Lego-Synthetic-Data
Dataset Summary
Paper | Github | HF Collection
SWE-Lego-Synthetic-Data contains 11.5k synthetic github issues (Python language) and their multi-turn agent trajectories. The column named messages is collected using Qwen/Qwen3-Coder-480B-A35B-Instruct with OpenHands (v0.53.0) agent scaffolding, which can be directly used for SFT training.
This dataset is part of the work presented in SWE-Lego, a supervised fine-tuning (SFT) recipe designed to achieve state-of-the-art performance… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/SWE-Lego-Synthetic-Data.hiring-bias-mitigation-synthetic-data
Hiring-bias mitigation — synthetic training data
Semi-synthetic data for training LLMs to make hiring decisions that do not depend on a
protected attribute (military status, gender, religion), in English and Ukrainian.
Real inputs, synthetic labels. CVs and job descriptions are real, anonymised postings
from the Djinni Recruitment Dataset (MIT). Decisions and rationales were written by the
teacher model Qwen/Qwen3.5-122B-A10B-GPTQ-Int4.
Code and results:… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/hiring-bias-mitigation-synthetic-data.Synthetic-Dataset-Childrens-Stories**Status: released 13-09-2026, repacked 14-09-2026.** The 14-09-2026 repack replaced 58
items after the acceptance gates were strengthened (prompt-instruction leaks, markdown
bullet lists and blockquotes); the other 29,942 are unchanged. Development stopped, pipeline released 17/09/26.
SAMPLE RELEASE: 30,000 synthetic children's short stories for early-reader language modelling.
Metrics
Value
genres
26
stories per genre
1.153-1.154K stories
total characters
38… See the full description on the dataset page: https://huggingface.co/datasets/ContextReq/Synthetic-Dataset-Childrens-Stories.Synthetic-JP-EN-Coding-Dataset-801k
Synthetic-JP-EN-Coding-Dataset-801k
Magpieによって作成したコードSFTデータセットであるAratako/Synthetic-JP-EN-Coding-Dataset-Magpie-69kを元に、Evol-Instructのような手法を用いて複数のinstructionとresonseを生成し拡張して作成した、日英混合801262件のコードSFT用合成データセットです。
日本語: 173849件
英語: 627413件
元のinstructionの作成に利用したモデルは以下の通りです。modelキーに該当レコードの作成に利用したモデル情報があります。
nvidia/Nemotron-4-340B-Instruct
microsoft/Phi-3-medium-4k-instruct
mistralai/Mixtral-8x22B-Instruct-v0.1… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-JP-EN-Coding-Dataset-801k.synthetic_vc_financial_decisions_reasoning_dataset
Best Curator Use Case in the Reasoning Datasets Competition: https://www.linkedin.com/feed/update/urn:li:activity:7330998995990781952/
Synthetic VC Financial Decisions Reasoning Dataset
Dataset Summary
The Synthetic VC Financial Decisions Reasoning Dataset is a large-scale collection designed to train, evaluate, and fine-tune language models on subjective, abstract financial reasoning tasks. It simulates venture capital (VC) workflows by capturing multiple… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/synthetic_vc_financial_decisions_reasoning_dataset.synthetic-swift-data-single-turn
Dataset Card for synthetic-swift-data-single-turn
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/vinhnx90/synthetic-swift-data-single-turn/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/vinhnx90/synthetic-swift-data-single-turn.GNOTHEIA-synthetic-insurance-dataset
GNOTHEIA Synthetic Insurance Dataset
Published by: Gratex International a.s.Project: InnovAIte — InnovAIte Slovakia
License: Apache 2.0Version: 1.0.0Contact: info@gratex.com
A synthetic insurance claims dataset designed for AI systems that evaluate insurance claims using OMG SBVR business rules, structured claim polycontexts and synthetic claim-related documents.
The dataset main goal is to support:
LLM fine-tuning pipeline
SBVR reasoning benchmarks
insurance claim AI… See the full description on the dataset page: https://huggingface.co/datasets/gratex/GNOTHEIA-synthetic-insurance-dataset.Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k
Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k
概要
5種類のオープンモデルとQwen/Qwen2.5-72B-Instruct-GPTQ-Int8を使って作成した、190854件の日本語合成Preferenceデータセットです。
以下、データセットの詳細です。
instructionには、Aratako/Magpie-Tanuki-8B-annotated-96kのinput_qualityがexcellentのものを利用
回答生成には、以下の5つのApache 2.0ライセンスのモデルを利用
weblab-GENIAC/Tanuki-8B-dpo-v1.0
team-hatakeyama-phase2/Tanuki-8x8B-dpo-v1.0-GPTQ-8bit
cyberagent/calm3-22b-chat
llm-jp/llm-jp-3-13b-instruct
Qwen/Qwen2.5-32B-Instruct-GPTQ-Int8… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k.database-query-logs-synthetic
Database Query Logs (synthetic)
3,995 database query-log entries spanning 10 engines - MySQL, PostgreSQL, MongoDB, SQL
Server, Oracle, MariaDB, SQLite, Cassandra, Redis, and Elasticsearch - with query text,
type, complexity, execution timing, and row-count metadata.
These queries are synthetic
The queries were programmatically generated, not captured from production systems.
They were produced by templating a set of query shapes across industry-flavored schema… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/database-query-logs-synthetic.synthetic-text2sql-dataset
Dataset Card for "synthetic-text2sql-dataset"
Dataset Summary
The synthetic-text2sql-dataset is a large-scale, structured dataset containing 100,000 training and 5,851 test examples designed to support research and development in SQL semantic parsing, text-to-SQL generation, and chain-of-thought (CoT) reasoning.
It was derived from an original DataFrame and converted into Hugging Face's datasets.Dataset format. Three new fields were added:
question: alias for the… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/synthetic-text2sql-dataset.Synthetic-JP-EN-Translation-Dataset-Magpie-Nemotron-4-20k
Synthetic-JP-EN-Translation-Dataset-Magpie-Nemotron-4-20k
Magpieの手法をnvidia/Nemotron-4-340B-Instructに対して適用し作成した、20000件の日⇔英翻訳データセットです。
データセットの作成にはDeepInfraを利用しました。
また、このリポジトリでデータセット作成に用いたコードを公開しています。これをベースに、システムプロンプトとstopを一部変更することで生成しています。
特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。
Synthetic-JP-Coding-Dataset-Magpie-Nemotron-4-10k
Synthetic-JP-Coding-Dataset-Magpie-Nemotron-4-10k
Magpieの手法をnvidia/Nemotron-4-340B-Instructに対して適用し作成した、約10000件の日本語のコーディング用対話データセットです。
データセットの作成にはDeepInfraを利用しました。
また、このリポジトリでデータセット作成に用いたコードを公開しています。これをベースに、システムプロンプトとstopを一部変更することで生成しています。
特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。
Synthetic-JP-EN-Coding-Dataset-Magpie-69k
Synthetic-JP-EN-Coding-Dataset-Magpie-69k
Magpieの手法を様々なモデルに対して適用し作成した、約69000件の日本語・英語のコーディング対話データセットです。
作成に利用したモデルは以下の通りです。modelキーに該当レコードの作成に利用したモデル情報があります。
nvidia/Nemotron-4-340B-Instruct
microsoft/Phi-3-medium-4k-instruct
mistralai/Mixtral-8x22B-Instruct-v0.1
cyberagent/calm3-22b-chat
データセットの作成にはDeepInfraを利用しました。
また、このリポジトリでデータセット作成に用いたコードを公開しています。これをベースに、プロンプトテンプレートやシステムプロンプト等を一部変更することで生成しています。特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。
building-engineering-synthetic-dataset-v5
Building Engineering Synthetic Dataset (V5)
Repository: Irfanuruchi/building-engineering-synthetic-dataset-v5
This repository contains a synthetic dataset for training engineering reasoning models focused on building engineering calculations and sanity checks.
The dataset was generated using physics-based engineering equations and structured prompts suitable for LLM fine-tuning.
It was used to train:
Irfanuruchi/qwen2.5-1.5b-buildeng-precheck-lora-v5
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/Irfanuruchi/building-engineering-synthetic-dataset-v5.Synthetic-Hinglish-Finetuning-Dataset
Hinglish Conversations Dataset
Overview
This dataset contains synthetically generated conversational dialogues in Hinglish (a blend of Hindi and English). The conversations revolve around typical college life, cultural festivities, daily routines, and general discussions, designed to be relatable and engaging.
Dataset Details
Language: Hinglish (Hindi + English)
Domain: College life, daily interactions, cultural events, and general discussions
Size: 3576… See the full description on the dataset page: https://huggingface.co/datasets/prakharb01/Synthetic-Hinglish-Finetuning-Dataset.synthetic_dataset_low-mid
Synthetic Dataset: Low Context, Medium Generation
Dataset Description
This is a synthetic benchmark dataset designed to test LLM inference performance in low-context, mid-generation scenarios. The dataset consists of 2,000 samples with randomly generated tokens that simulate workloads where models receive short prompts but generate longer responses.
Use Cases
This dataset is ideal for benchmarking:
Creative writing and content generation
Code generation from… See the full description on the dataset page: https://huggingface.co/datasets/jonasluehrs-jaai/synthetic_dataset_low-mid.synthetic-medical-mistakes-dataset
Synthetic Medical Mistakes Dataset (SFT Training Data)
A dataset of 350 synthetic clinical reports with gold-standard error annotations, generated by state-of-the-art LLMs for supervised fine-tuning of clinical error detection models. Created as part of the Clinipal project.
Dataset Description
Overview
This dataset was designed to train AI models to detect critical patient safety errors in clinical documentation. Each entry contains a synthetic emergency… See the full description on the dataset page: https://huggingface.co/datasets/Vrda/synthetic-medical-mistakes-dataset.synthetic-data-papers
Synthetic Data Papers — FineSet
A research-paper dataset on Synthetic Data Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-12.
It is not auto-updated. Research on Synthetic Data Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this dataset
Quality-scored: quality_score float… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/synthetic-data-papers.synthetic_dataset_mid-mid
Synthetic Dataset: Medium Context, Medium Generation
Dataset Description
This is a synthetic benchmark dataset designed to test LLM inference performance in balanced workload scenarios. The dataset consists of 2,000 samples with randomly generated tokens that simulate realistic, mixed-workload patterns where both input and output lengths are moderate.
Use Cases
This dataset is ideal for benchmarking:
Chat applications with conversational exchanges
API services… See the full description on the dataset page: https://huggingface.co/datasets/jonasluehrs-jaai/synthetic_dataset_mid-mid.Norwegian-Synthetic-HR-data-v-1
Synthetic norwegian public sector HR dataset
Dataset description
This dataset contains 4,000 rows of synthetic instructional data focused on Human Resources (HR) topics within the Norwegian public sector.
The license for the dataset follows the license of the LLMs used to generate the data. Users are advised to review the specific terms associated with the source models before use.
The datasets includes Chain of Thought (CoT) reasoning traces and is generated using a… See the full description on the dataset page: https://huggingface.co/datasets/Hebbelille/Norwegian-Synthetic-HR-data-v-1.synthetic-patient-dr-data
Synthetic Patient DR Data
Synthetic doctor-patient consultation dataset with structured clinical outputs and optional full-consultation audio.
Dataset Summary
This dataset was generated for research and prototyping in:
clinical dialogue generation
structured clinical extraction
text-to-audio workflows
conversational healthcare modeling
All consultations are synthetic and should not be treated as real clinical encounters.
Export Metadata
Mode: audio
Repo… See the full description on the dataset page: https://huggingface.co/datasets/TumeloKonaite/synthetic-patient-dr-data.synthetic-irc-data
Synthetic IRC Conversation Dataset
Dataset Description
This dataset contains 1,500 synthetic IRC-style conversations featuring multiple participants, including an AI character named Em. The conversations were generated to replicate authentic IRC chat dynamics with natural flow, interruptions, and varied engagement levels.
Dataset Summary
Total conversations: 1,500
Total size: ~10MB
Format: JSONL with IRC-style formatting
Language: English
License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/david-ar/synthetic-irc-data.instruction-dataset-qwen-synthetic
Synthetic Instruction Dataset
A high-quality synthetic instruction-response dataset generated using Qwen 2.5 72B Instruct.
Dataset Details
Total Examples: 500
Train/Test Split: 475/25
Generator Model: Qwen 2.5 72B Instruct
Generated: 2026-05-30
Categories
The dataset covers diverse categories:
Coding & Programming - Algorithm implementation, code explanation, debugging
Mathematics & Reasoning - Problem solving with step-by-step explanations… See the full description on the dataset page: https://huggingface.co/datasets/himanshunakrani9/instruction-dataset-qwen-synthetic.matchcv-synthetic-dataset
MatchCV Synthetic Training Dataset
500 synthetic, anonymous CV/Job matching pairs — generated algorithmically, zero real personal data.
Used to fine-tune sallani/MatchCV-Qwen2.5-0.5B.
What's inside
Each record is an instruction-following example:
{
"instruction": "Analyse le profil candidat...",
"input": "=== CV ANONYMISÉ ===\n...\n=== OFFRE ===\n...",
"output": "Score de matching : 73.5/100\nCompétences correspondantes : ...\nGaps techniques :… See the full description on the dataset page: https://huggingface.co/datasets/sallani/matchcv-synthetic-dataset.myX-Burmese-Synthetic-Pseudo-Syllables
📝 Burmese Synthetic Pseudo-Syllables Dataset
This dataset contains 5,814,699 computer-generated (synthetic) Myanmar pseudo-syllables structured systematically based on specific complex linguistic and orthographic patterns.
Developed as part of the foundational research for low-resource language processing, this dataset serves as a rigorous baseline and stress-testing environment for Burmese Natural Language Processing (NLP), tokenization, font rendering, and spell-checking… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myX-Burmese-Synthetic-Pseudo-Syllables.synthetic-rag-dataset
Dataset Card for synthetic-rag-dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/m-newhauser/synthetic-rag-dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/m-newhauser/synthetic-rag-dataset.ADHD-Synthetic-Dataset
ADHD Assistant Synthetic Chat Dataset (JSONL - ChatML)
📌 Latar Belakang & Urgensi Data
Dataset ini dibangun karena absennya dataset open-source yang didedikasikan untuk melatih model bahasa (LLM/SLM) dalam memahami dan membantu kondisi neurodivergen, khususnya ADHD. Menjadi berbeda di tengah lingkungan yang kaku adalah proses yang berat bagi mereka yang belum mampu menerima dirinya sendiri. Dataset ini dirancang dengan penuh kehati-hatian untuk mengisi kekosongan… See the full description on the dataset page: https://huggingface.co/datasets/kareem2808/ADHD-Synthetic-Dataset.Synthetic_CoT_dataset_RUСинтетический русский датасет для Chain-of-Thought (CoT) представляет собой набор текстов, созданных для тренировки моделей в пошаговом рассуждении. Каждый элемент включает входной запрос, последовательность промежуточных шагов рассуждения и окончательный ответ. Цель датасета – улучшить способность моделей формировать логические объяснения и решения сложных задач. Примеры задач охватывают арифметические вычисления, вопросы по общей эрудиции, логические и аналитические задачи. Данные… See the full description on the dataset page: https://huggingface.co/datasets/DataSynGen/Synthetic_CoT_dataset_RU.synthetic-data-factory-5000-20260317
Synthetic Data Factory 5k
This dataset contains 5,000 synthetic math and logic examples generated without LLM-based generation.
File
generated_5000.jsonl: JSONL rows with problem text, explanation text, final answer, tags, metadata, and quality report.
Families
arithmetic expression evaluation
linear equation solving
comparison logic / transitive reasoning
Generation approach
Examples are created through a world-model-first pipeline:
latent… See the full description on the dataset page: https://huggingface.co/datasets/Abhiram1009/synthetic-data-factory-5000-20260317.gk-synthetic-data-2026-ko
gk-synthetic-data-2026-ko.jsonl
gk-synthetic-data-2026-ko.jsonl은 외부 teacher 모델의 고급 지식과 장문 설명 능력을 한국어 SFT용으로 증류한 데이터셋이다.
현재 행 수: 77,408
한 줄 용도: 2026년 기준 고급 지식과 장문 설명 능력을 보강하기 위한 한국어 합성 SFT 데이터셋
주 용도: kanana 기반 지식/추론/장문 응답 모델 SFT
형식: JSONL, 한 줄에 하나의 학습 예시
스키마: {"messages":[{"role":"user","content":"..."},{"role":"assistant","content":"..."}]}
일부 행은 assistant 답변의 첫 부분에 공개 판단 근거 형식의 <think>...</think> 블록을 포함한다. 현재 포함 행 수는 111개이며, 닫는 태그 뒤에는 장문 최종 답변이 이어진다.
데이터 의미… See the full description on the dataset page: https://huggingface.co/datasets/Algocean/gk-synthetic-data-2026-ko.
