datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.synthetic_multilingual_llm_prompts
Image generated by DALL-E. See prompt for more details
📝🌐 Synthetic Multilingual LLM Prompts
Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.synthetic-legal
⚖️ Synthetic Legal (Query, Response) Dataset
📚 140,000 synthetic (legal query, legal response) pairs across 13 legal domains, built to resemble the structure of real-world fact patterns and citation-backed answers.
⚠️ Disclaimer: All text is synthetically generated and IS NOT LEGALLY ACCURATE. Citations are real but assigned at random, and the verified_solution and verification_method columns are template labels, not evidence of review. This dataset is not legal advice.… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/synthetic-legal.cyberattack-blockchain-synth
ELTEX-Blockchain: A Domain-Specific Dataset for Cybersecurity
🔐 12k Synthetic Social Media Messages for Early Cyberattack Detection on Blockchain
Dataset Statistics
Category
Samples
Description
Cyberattack
6,941
Early warning signals and indicators of cyberattacks
General
4,507
Regular blockchain discussions (non-security related)
Dataset Structure
Each entry in the dataset contains:
message_id: Unique identifier for each message… See the full description on the dataset page: https://huggingface.co/datasets/dn-institute/cyberattack-blockchain-synth.synthetic-fine-arts
🎨 Synthetic Fine Arts (Challenge, Solution) Dataset
🖼️ 225,000 synthetic (artistic challenge, proposed solution) pairs spanning Visual Arts, Performing Arts, Musical Arts, Literary Arts, Digital Arts, Art History, and Art Theory, with metadata fields that mirror the shape of a curation workflow.
⚠️ Disclaimer: All text is synthetically generated and should not be relied on for artistic, historical, or technical accuracy. The VerificationMethod, ReferenceMaterial, and… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/synthetic-fine-arts.MedSyn-synthetic
Synthetic dataset:
generated.csv - synthetic datasets containing 41,185 clinical note samples spanning 219 ICD-10 codes.
Data field
Description
idx
Unique sample identifier.
ICD-10
The targeted ICD-10 code used for prior data sampling.
generation_model
The model used for sample generation (GTP-3.5, GPT-4, LLaMA-7b, LLaMA-13b)
prompt
Prompt used for sample generation.
prior
Type of prior data used for sample generation.
example
Bool variable for the presence or… See the full description on the dataset page: https://huggingface.co/datasets/Glebkaa/MedSyn-synthetic.synthetic_polistance
Fully Synthetic Prompts for LLM Political Stance Detection
All resources developed in the article "Templated or fully Synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance" (Chalkidis, 2026).
Paper Abstract
Political stance detection in LLMs has long been dominated by closed-ended, multiple-choice political survey questions—originally designed for humans, and thus lacks the realism and nuance of human-AI… See the full description on the dataset page: https://huggingface.co/datasets/kiddothe2b/synthetic_polistance.synthetic_discharge_summ
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is a subset of the dataset for Asclepius model (arxiv).
The original dataset is made up of synthetic notes generated from PMC-Patients case reports with GPT-3.5.
We filtered the summarization task for discharge notes. The dataset contains 13,584 notes.
Supported Tasks
This dataset covers below summarization task
Languages
English
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/bluesky333/synthetic_discharge_summ.SynPerForm-synthetic-persian-formality-pairs
SynPerForm
SynPerForm is a paired Persian dataset for formality style
transfer. Each informal text is paired with a freely written formal rewrite
that preserves its meaning without requiring lexical or structural
equivalence. The formal rewrites were generated with OpenAI GPT-5.6 Luna.
Columns
Informal: the original informal Persian text.
Formal: a free formal rewrite that preserves the original meaning.
Intended uses
Formality style transfer:… See the full description on the dataset page: https://huggingface.co/datasets/shekar-ai/SynPerForm-synthetic-persian-formality-pairs.indic-synthetic-profiles
🇮🇳 Indian Synthetic Identity Dataset
10,000 realistic Indian synthetic identities across 8 languages — generated by indic-faker
Dataset Description
This dataset contains 10,000 rows of realistic, synthetic Indian identity data generated using the indic-faker Python library. Every record is algorithmically valid — Aadhaar numbers pass Verhoeff checksum verification, GSTINs have correct state codes, and names are culturally authentic across 8 Indian languages.… See the full description on the dataset page: https://huggingface.co/datasets/adwaith06/indic-synthetic-profiles.synthanime-openhermes2.5
What is this dataset?
This is a hybrid Instruction-Response dataset (scraped + synthetic) for anime synopses.
Given around 10,000 scraped anime synopses, the user instructions for this dataset were generated using Teknium's Openhermes 2.5
The assistant column consists of synopsis for different animes (which were previously scraped)
The user column consists of the instructions that "might have" generated the synopsis (synthetically generated)
The goal of this dataset is: to be used… See the full description on the dataset page: https://huggingface.co/datasets/maharshipandya/synthanime-openhermes2.5.Synthetic-Context-Conversations
Synthetic-Context-Conversations
Overview
The Synthetic-Context-Conversations dataset is a collection of synthetic conversations designed to simulate empathetic and context-rich dialogues. It is particularly useful for tasks such as text generation, summarization, and question answering. The dataset is available in English and contains between 10,000 to 100,000 entries.
Dataset Details
Modalities: Text
Languages: English
Size: 10K-100K
Formats: Parquet
License:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Synthetic-Context-Conversations.BG-GSM-Synthetic
Overview
This dataset aims to create a bulgarian version of the GSM-Symbolic dataset
Code to generate samples is available at our Github
BG-GSM-Symbolic
Synthetic Dataset Generation for Large Language Model evaluation. The dataset is generated in the style of GSM-8k, thus it contains relatively easy mathematical questions that a bright grade school student should be able to answer. This project is intended for automatic evaluation of LLMs, which have already been trained… See the full description on the dataset page: https://huggingface.co/datasets/WT-solutions/BG-GSM-Synthetic.Referencing_Errors_Synthetic_EN
Synthetic Dataset for Automatic Error Correction in Referencing
This dataset includes 4,600 parallel sentences for Automatic Error Correction in referencing in German.It was synthetically created with gpt-4o-mini model according to the Institutional Guidelines of the Center for Translation Studies (CTS), University of Vienna.
Dataset Description
corrupted_sentence: the sentence containing the referencing error
clean_sentence: the correct version of the corrupted sentence… See the full description on the dataset page: https://huggingface.co/datasets/elizaveta-dev/Referencing_Errors_Synthetic_EN.Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/Xavier1234/Asclepius-Synthetic-Clinical-Notes.arabic-mantoq-synthetic-g2pcyberattack-blockchain-synth
ELTEX-Blockchain: A Domain-Specific Dataset for Cybersecurity
🔐 12k Synthetic Social Media Messages for Early Cyberattack Detection on Blockchain
Dataset Statistics
Category
Samples
Description
Cyberattack
6,941
Early warning signals and indicators of cyberattacks
General
4,507
Regular blockchain discussions (non-security related)
Dataset Structure
Each entry in the dataset contains:
message_id: Unique identifier for each message… See the full description on the dataset page: https://huggingface.co/datasets/cmp81/cyberattack-blockchain-synth.synthetic-thai-follow-up-question3D_Synthetic_Petroleum_Derived_GEMS
3D Synthetic Petroleum-Derived GEMS (MVP Release)
📌 Dataset Overview
This dataset contains 197 elite, highly complex 3D molecular structures derived from petroleum fractions. Designed specifically for petrochemicals, materials science, organic semiconductors, and specialized additives, these compounds represent a curated "Golden Fund" of stable, complex hydrocarbons.
Unlike drug-like molecules, this dataset focuses on polycyclic architectures, rigid 3-ring… See the full description on the dataset page: https://huggingface.co/datasets/nadizik/3D_Synthetic_Petroleum_Derived_GEMS.Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/LampsteR/Asclepius-Synthetic-Clinical-Notes.tweepfake_synthetic
TweepFake Synthetic Tweets
Synthetic tweets generated by three open / API LLMs, conditioned on topics extracted from
real human tweets. Built as a controlled-variable corpus for studying — and detecting —
machine-generated social-media text.
The design holds the topic constant across conditions: every topic is fed to multiple
models under multiple prompting strategies, so differences in the output can be attributed to
the model and prompt, not the subject matter.… See the full description on the dataset page: https://huggingface.co/datasets/KirillNik/tweepfake_synthetic.hopepet-ai-synthetic-dataset
HOPEPET AI Synthetic Pet-Care Dataset
This dataset was created for the HOPEPET AI final project.
HOPEPET AI is an AI-based pet-care prototype for dog and cat owners.The goal of the system is to provide responsible first-step guidance when a user notices a pet-related problem, symptom, or behavior change.
The system does not provide veterinary diagnosis, does not replace a veterinarian, and does not recommend medication or dosages.
Dataset Description
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/shak9345/hopepet-ai-synthetic-dataset.portuguese-sentences-synthetic-g2p
Dataset Card for 'TigreGotico/portuguese_g2p'
Dataset Description
Dataset Summary
TigreGotico/portuguese_g2p is a Grapheme-to-Phoneme (G2P) dataset for Portuguese, offering phonetic transcriptions for sentences across ten different regional variants.
It is derived from the portuguese_phonetic_lexicon and is designed to aid in the development of robust Speech Recognition (ASR) and Text-to-Speech (TTS) models that account for dialectal variation in Portuguese.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-sentences-synthetic-g2p.Referencing_Errors_Synthetic_DE
Synthetic Dataset for Automatic Error Correction in Referencing
This dataset includes 4,600 parallel sentences for Automatic Error Correction in referencing in German.It was synthetically created with gpt-4o-mini model according to the Institutional Guidelines of the Center for Translation Studies (CTS), University of Vienna.
Dataset Description
corrupted_sentence: the sentence containing the referencing error
clean_sentence: the correct version of the corrupted sentence… See the full description on the dataset page: https://huggingface.co/datasets/elizaveta-dev/Referencing_Errors_Synthetic_DE.synthetic-knowledge-itemssynth-dent-test
