datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-Terminal-Synthetic-Tasks
Terminal-Corpus: Task Structure Specification
This repository contains the skill-based synthetic tasks within the Terminal-Corpus. These tasks are designed to evaluate and train autonomous agents in realistic Linux terminal environments.
🏗️ Task Anatomy
Each task is contained within a dedicated directory and follows a strict four-component architecture:
1. Instruction (instruction.md)
Purpose: Provides the natural language description of the objective.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Terminal-Synthetic-Tasks.context_qa_sum_qwen3_synthetic
Context-based QA and Summarization Synthetic Dataset
Overview
This dataset contains synthetic context-based question-answering (QA) and summarization data. The data was synthesized using:
Source context: openbmb/Ultra-FineWeb
Synthesis model: Qwen3-30B-A3B-Instruct-2507
Each context is obtained by taking the initial segment of raw pretraining text from Ultra-FineWeb, truncated to at most the corresponding number of tokens, while ensuring the truncation does not occur in… See the full description on the dataset page: https://huggingface.co/datasets/yuyijiong/context_qa_sum_qwen3_synthetic.synthetic_text_to_sql
Image generated by DALL-E. See prompt for more details
synthetic_text_to_sql
gretelai/synthetic_text_to_sql is a rich dataset of high quality synthetic Text-to-SQL samples,
designed and generated using Gretel Navigator, and released under Apache 2.0.
Please see our release blogpost for more details.
The dataset includes:
105,851 records partitioned into 100,000 train and 5,851 test records
~23M total tokens, including ~12M SQL tokens
Coverage across 100 distinct… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_text_to_sql.qa_squadshifts_synthetic
Dataset Card for "lmqg/qa_squadshifts_synthetic"
Dataset Summary
This is a synthetic QA dataset generated with fine-tuned QG models over lmqg/qa_squadshifts, made for question-answering based evaluation (QAE) for question generation model proposed by Zhang and Bansal, 2019.
The test split is the original validation set of lmqg/qa_squadshifts, where the model should be evaluate on.
Supported Tasks and Leaderboards
question-answering
Languages… See the full description on the dataset page: https://huggingface.co/datasets/lmqg/qa_squadshifts_synthetic.Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.Synthetic-AI-ML-Dataset
Synthetic-AI-ML-Dataset
Synthetic Q&A dataset on AI and Machine Learning
Dataset Details
Metric
Value
Topic
AI and Machine Learning
Total Q&A Pairs
14021
Valid Pairs
14021
Provider/Model
ollama/gpt-oss:120b
Generation Cost
Metric
Value
Prompt Tokens
14,941,957
Completion Tokens
17,159,263
Total Tokens
32,101,220
GPU Energy
12.9628 kWh
Sources
This dataset was generated from 474 scholarly papers:
#… See the full description on the dataset page: https://huggingface.co/datasets/Kylan12/Synthetic-AI-ML-Dataset.gretel-synthetic-text-to-sql
Fork of gretelai/synthetic_text_to_sql
The gretelai/synthetic_text_to_sql dataset is a large, Apache 2.0 licensed, synthetic Text-to-SQL dataset consisting of 105,851 high-quality records across 100 diverse domains, designed for training language models. It includes comprehensive SQL tasks with varying complexities, database contexts, natural language explanations, and contextual tags, outperforming existing datasets in SQL correctness and standards compliance.
synthetic-pre1930-sftTL;DR
A vintage finetuning dataset (~416k rows, eleven task routes). Sourced by taking excerpts
from pre-1930's texts, turning these into verbatim answers, and then using deepseek-chat to
generate period-appropriate questions of those answers. Any model tuned on this dataset should,
theoretically, never update its weights on anachronistic text, since questions are masked in the
finetuning stages. Features composition, verse, narrative, reasoning, multiturn dialogue, and
calibrated uncertainty… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/synthetic-pre1930-sft.gsm8k-synthetic-diverse-8b
gretelai/gsm8k-synthetic-diverse-8b
This dataset is a synthetically generated version inspired by the GSM8K https://huggingface.co/datasets/openai/gsm8k dataset, created entirely using Gretel Navigator with meta-llama/Meta-Llama-3.1-8B as the agent LLM. It contains ~1500 Grade School-level math word problems with step-by-step solutions, focusing on age group, difficulty, and domain diversity.
Key Features:
Synthetically Generated: Math problems created using Gretel… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gsm8k-synthetic-diverse-8b.Retrieval-Synthetic-NVDocs-v1
Dataset Description:
Retrieval-Synthetic-NVDocs-v1 is a synthetic retrieval dataset with question–answer supervision designed to train and evaluate embedding and RAG systems. The dataset was generated on top of NVIDIA's publicly available content using NeMo Data Designer, NVIDIA's open-source framework for generating high-quality synthetic data from scratch or based on seed data.
The dataset contains document chunks paired with semantically rich question-answer pairs across multiple… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Retrieval-Synthetic-NVDocs-v1.tw-legal-synthetic-qa
Dataset Card for tw-legal-synthetic-qa
Dataset Summary
本合成對話資料集(下稱本資料集)由 THUDM/chatglm3-6b-32k 和 lianghsun/tw-processed-judgments,由實驗後的 prompt 去生成繁體中文法律對話合成集。
Supported Tasks and Leaderboards
本資料集可以運用在 SFT,讓模型學會如何回答法律問題。
Languages
繁體中文。
Dataset Structure
Data Instances
一個資料樣本如下,首先由 user 發問了一個具有(或可能有)法律情境的問題,然後 assistant 回答法律相關知識。
{
"messages":[
{
"role":"user"… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-synthetic-qa.synthetic-stroke-patients-v3
Synthetic Stroke Patient Episodes v3
Release status: PILOT_ONLY. Clinical use is not approved.
This repository contains 1,000 wholly synthetic adult stroke episodes with 130 structured dimensions, compositional clinical notes, and 5,000 time-indexed educational task examples. It contains no real patient records.
Splits
Train: 700 patients
Validation: 100 patients
Test: 100 patients
Challenge: 100 patients
Each episode is also represented at five information… See the full description on the dataset page: https://huggingface.co/datasets/therayelab/synthetic-stroke-patients-v3.synthetic-it-support-tickets
Synthetic IT Support Tickets — PII-Enriched + Redaction Ground Truth
745 synthetic IT service-management incident records for LLM wiki and
retrieval-augmented-generation experiments. Each record is a help-desk/IT-ops incident with
submitted ticket text, timestamped troubleshooting correspondence, structured diagnostics, root
cause, and resolution steps.
The free text is enriched with realistic technical detail and injected synthetic PII. The corpus
ships two authored… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthetic-it-support-tickets.Instella-GSM8K-synthetic
Instella-GSM8K-synthetic
The Instella-GSM8K-synthetic dataset was used in the second stage pre-training of Instella-3B model, which was trained on top of the Instella-3B-Stage1 model.
This synthetic dataset was generated using the training set of GSM8k dataset, where we first used Qwen2.5-72B-Instruct to
Abstract numerical values as function parameters and generate a Python program to solve the math question.
Identify and replace numerical values in the existing question with… See the full description on the dataset page: https://huggingface.co/datasets/amd/Instella-GSM8K-synthetic.VisOnlyQA_Eval_Synthetic
VisOnlyQA
🌐 Project Website | 📄 Paper | 🤗 Dataset | 🔥 VLMEvalKit
This repository contains the code and data for the paper "VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information" (COLM 2025).
VisOnlyQA is designed to evaluate the visual perception capability of large vision language models (LVLMs) on geometric information of scientific figures. The evaluation set includes 1,200 mlutiple choice questions in 12 visual perception… See the full description on the dataset page: https://huggingface.co/datasets/ryokamoi/VisOnlyQA_Eval_Synthetic.synthetic_vc_financial_decisions_reasoning_dataset
Best Curator Use Case in the Reasoning Datasets Competition: https://www.linkedin.com/feed/update/urn:li:activity:7330998995990781952/
Synthetic VC Financial Decisions Reasoning Dataset
Dataset Summary
The Synthetic VC Financial Decisions Reasoning Dataset is a large-scale collection designed to train, evaluate, and fine-tune language models on subjective, abstract financial reasoning tasks. It simulates venture capital (VC) workflows by capturing multiple… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/synthetic_vc_financial_decisions_reasoning_dataset.ww2-synthetic-corpus
WWII Synthetic LLM Training Corpus
A large synthetic dataset of conversational training examples about World War II, generated by a custom synthetic-data pipeline. The knowledge originates in curated Wikipedia articles; large language models were used only as transformation tools to reformat that source knowledge into diverse training patterns — they are not the source of the facts.
Total examples: 16,397,795
Subsets (configs): 92
Format: ChatML-style messages (role/content)… See the full description on the dataset page: https://huggingface.co/datasets/wayneworkman2012/ww2-synthetic-corpus.synthetic-swift-data-single-turn
Dataset Card for synthetic-swift-data-single-turn
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/vinhnx90/synthetic-swift-data-single-turn/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/vinhnx90/synthetic-swift-data-single-turn.GNOTHEIA-synthetic-insurance-dataset
GNOTHEIA Synthetic Insurance Dataset
Published by: Gratex International a.s.Project: InnovAIte — InnovAIte Slovakia
License: Apache 2.0Version: 1.0.0Contact: info@gratex.com
A synthetic insurance claims dataset designed for AI systems that evaluate insurance claims using OMG SBVR business rules, structured claim polycontexts and synthetic claim-related documents.
The dataset main goal is to support:
LLM fine-tuning pipeline
SBVR reasoning benchmarks
insurance claim AI… See the full description on the dataset page: https://huggingface.co/datasets/gratex/GNOTHEIA-synthetic-insurance-dataset.japanese-math-synthetic-108k
clean_v1_noassert_allcats
LLM-JP チューニングコンテスト 2026 向けに作成した日本語数学問題の学習データセットです。
中学1年〜高校数学IIICまでの各カテゴリを網羅しています。問題作成・解法生成・検証のすべてを GPT-OSS (120B) で行い、重複除去を経てクリーニング済みです。
データ概要
2種類の分類軸でデータを収録しています。
学年別(中学〜高校カリキュラム準拠)
ファイル
カテゴリ
行数
主な単元
chu1.jsonl
中学1年
9,610
正負の数, 文字式, 一次方程式, 比例反比例
chu2.jsonl
中学2年
9,761
文字式, 一次関数, 連立方程式, 確率
chu3.jsonl
中学3年
9,454
二次方程式, 二次関数, 平方根, 展開と因数分解
IA.jsonl
数学IA
9,881
整数の性質, 場合の数と確率, 2次関数, 数と式
IIB.jsonl
数学IIB
11,089
数列, いろいろな式… See the full description on the dataset page: https://huggingface.co/datasets/yamaTK/japanese-math-synthetic-108k.synthetic_multilingual_llm_prompts
Image generated by DALL-E. See prompt for more details
📝🌐 Synthetic Multilingual LLM Prompts
Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.synthetic-gsm8k-evolutionary-405b
gretelai/synthetic-gsm8k-evolutionary-405b
This dataset is a synthetically generated version inspired by the GSM8K dataset, created entirely using Gretel Navigator with meta-llama/Meta-Llama-3.1-405B as the agent LLM. It contains Grade School-level reasoning tasks with step-by-step solutions, focusing on multi-step reasoning problems.
Key Features:
Synthetically Generated: Built using Gretel Navigator, leveraging evolutionary approach for diversity to create both the… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic-gsm8k-evolutionary-405b.LLM_Electrical_Engineering_Educational_Synthetic_DialogDataset Card for LLM_Electrical_Engineering_Educational_Synthetic_Dialog
The full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com
Want more? 🚀 Get the AI Startup Bundle from Gumroad.
Dataset Description
The LLM_Electrical_Engineering_Educational_Synthetic_Dialog dataset contains AI-generated conversational interactions designed for training large language models in electrical engineering education. This synthetic dialogue corpus simulates tutor-student… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/LLM_Electrical_Engineering_Educational_Synthetic_Dialog.nmr-dataset-synthetic-retraction
NMR Dataset: Synthetic Retraction
This dataset contains 360 complete belief-revision episodes: 90 each for
monotonic, nmr_new_evidence, nmr_retraction, and nmr_mixed.
Each JSONL row is one episode, with a fixed dependency graph, an initial
belief base, three revisions, and complete gold T/F/U states at all four
checkpoints. It is intentionally not expanded into one row per proposition.
An evaluator can construct a query for every proposition at any checkpoint
from the episode… See the full description on the dataset page: https://huggingface.co/datasets/leo-bjpark/nmr-dataset-synthetic-retraction.synthetic-superconductor-materials-dataset
synthetic-superconductor-materials-dataset
Synthetic Q&A dataset on Superconductor Materials, generated with SDGS (Synthetic Dataset Generation Suite).
Dataset Details
Metric
Value
Topic
Superconductor Materials
Total Q&A Pairs
2649
Valid Pairs
2649
Provider/Model
ollama/gpt-oss:120b
Sources
This dataset was generated from 170 scholarly papers:
#
Title
Authors
Year
Source
QA Pairs
1
Observation of a large-gap… See the full description on the dataset page: https://huggingface.co/datasets/Kylan12/synthetic-superconductor-materials-dataset.synthetic-legal
⚖️ Synthetic Legal (Query, Response) Dataset
📚 140,000 synthetic (legal query, legal response) pairs across 13 legal domains, built to resemble the structure of real-world fact patterns and citation-backed answers.
⚠️ Disclaimer: All text is synthetically generated and IS NOT LEGALLY ACCURATE. Citations are real but assigned at random, and the verified_solution and verification_method columns are template labels, not evidence of review. This dataset is not legal advice.… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/synthetic-legal.tiny-llm-synthetic-qa
Tiny-LLM: Synthetic Question-Answering Dataset
Dataset Description
This dataset was created for the fine-tuning stage of the Tiny-LLM Project, a project focused on training and evaluating compact language models from scratch.
It contains 706,727 high-quality, synthetic multi-turn Question-Answering (Q&A) conversations in English, generated using the Gemini API. The dataset was designed to teach small models instruction-following capabilities across a diverse range of… See the full description on the dataset page: https://huggingface.co/datasets/Gabriel8/tiny-llm-synthetic-qa.ViLegalQA-Synthetic-Curation
ViLegalQA Synthetic Curation
Dataset summary
This repository releases the synthetic Vietnamese legal QA research artifacts produced in the accompanying study. The primary resource contains 10,095 synthetic QA items spanning true/false, multiple-choice, and open-ended tasks. It is accompanied by the final curation/quality annotations used in the study, plus aggregated labels for 600 items from the five-expert human calibration panel.
Manuscript: Human-Calibrated… See the full description on the dataset page: https://huggingface.co/datasets/nguyenkhanh87/ViLegalQA-Synthetic-Curation.math-dow-mod-synthetic-v2
Cyclic Calendar Story Reasoning v2
v2 of liodon-ai/math-dow-mod-synthetic-v1,
rebuilt around one piece of feedback from running v1 at scale: story and
reasoning diversity has to scale to millions of rows, or it reads as one
template repeated. v1's story setups were picked off a fixed list of 8
full sentences per task/direction, and every reasoning trace followed one
fixed 4-slot skeleton (opener / op / mod-reduce / conclude) with only the
wording swapped — fine at 10-50k rows… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/math-dow-mod-synthetic-v2.math-dow-mod-synthetic-v1
Math + Cyclic-Time Synthetic Dataset
Synthetic dataset for training a small (~10M-100M param), task-specialized
LLM on arithmetic (addition, multiplication), cyclic time arithmetic
(days-of-week, months, 12-hour and 24-hour clock), and mod-k remainder
probes — generalization-focused rather than memorization, following on
from the T2/T5/T10 modular-circuit discussion. Days-of-week, months, and
hours are all instances of the same underlying cyclic/modular-addition
structure… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/math-dow-mod-synthetic-v1.
