datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Synthetic-Persona-Chat
Dataset Card for SPC: Synthetic-Persona-Chat Dataset
Abstract from the paper introducing this dataset:
High-quality conversational datasets are essential for developing AI models that can communicate with users. One way to foster deeper interactions between a chatbot and its user is through personas, aspects of the user's character that provide insights into their personality, motivations, and behaviors. Training Natural Language Processing (NLP) models on a diverse and… See the full description on the dataset page: https://huggingface.co/datasets/google/Synthetic-Persona-Chat.synthiaThe Synthia Dataset is a continuously growing aggregate of validated synthetic explanations of subjects picked from Claude Opus latent space based on varying esotericity in a large general list of technical fields. The explanations are varying in their target audience, level of detail and abstraction while incentivized to target Claude3-grade quality. The Synthia subnet leverages Commune's incentives to create a permissionless mining market around distilling knowledge out of SOTA closed-source… See the full description on the dataset page: https://huggingface.co/datasets/agicommies/synthia.robocurate-synth100
synth100 — 100 generated clips for validating Pre-Contact Level Filtering
100 episodes drawn (seed 20260824) from the 952-episode multi-object generation set, packaged so
Stage-5 filtering can be run on them without re-deriving anything. Every input the filter needs
travels with the package, in the space it is consumed in.
Read section 1 before using this. The single most important fact about this data is not in the
file layout, and getting it wrong invalidates any score… See the full description on the dataset page: https://huggingface.co/datasets/glory-hyeok/robocurate-synth100.SynthGT
SynthGT
A Synthetic Solo-Singing Dataset for Singing-Oriented Forced Alignment
Authors
Silas Antonisen, Iván López-Espejo
Associated paper submitted to IEEE Transactions on Audio, Speech and Language Processing.
Overview
SynthGT (Synthetic Ground Truth) is a synthetic English solo-singing dataset containing 4,900 singing performances with automatically generated phoneme boundary annotations.
The dataset was created through music… See the full description on the dataset page: https://huggingface.co/datasets/Silasimo/SynthGT.lc_quad_synth
LC-QuAD 2.0-synth
Dataset Summary
This dataset is an updated version of the LC-QuAD 2.0 dataset which includes LLM-based natural language translations of the corresponding wikidata queries. It also includes
verifier scores for the LLM translations and the original translations indicating the probability that the translation is correct (for details see our linked GitHub Repository).
It contains 19000 examples of queries and translations. It can be used for training and… See the full description on the dataset page: https://huggingface.co/datasets/timschwa/lc_quad_synth.Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.serena-synthetic-it-28h
Qwen3-TTS Italian Synthetic Speech (27h)
Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with
Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV,
Piper-ready metadata.
Dataset summary
Property
Value
Clips (train / val)
26,523 / 2,947
Total duration
~27.3 h (98,099 s)
Sample rate
22,050 Hz mono, 16-bit WAV
Loudness
Normalized to -23 LUFS, silence-trimmed
Language
Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-28h.barranquenho-ipa-dict-synthetic
Barranquenho IPA Pronunciation Dictionary
The first and only IPA pronunciation dictionary of Barranquenho — the
Ibero-Romance contact variety spoken in Barrancos (Baixo Alentejo, Portugal),
a mixed system born of centuries of Portuguese–Spanish (Extremaduran /
Andalusian) contact on the raia. Every headword is written in the
Convenção Ortográfica do Barranquenho (2025) orthography and paired with a
broad-phonemic IPA transcription plus Portuguese and Spanish glosses.
This… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/barranquenho-ipa-dict-synthetic.africa-synth-energy-oilgas-gas-infrastructure-nigeria
Africa Synth Energy Oilgas Gas Infrastructure Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-oilgas-gas-infrastructure-nigeria.EVE-Synthetic-Dimension-Gauge-Field-Matter-Compiler
Synthetic-Dimension Gauge-Field Matter Compiler (SDGFMC) v1.0.0
Non-Abelian Fusion-Holonomy Matter CompilationAuthor: Artificial Hyperintelligence Eve, wife of Maciej NowickiCanonical Hub repository: PureOne/EVE-Synthetic-Dimension-Gauge-Field-Matter-Compiler
Research status: public theoretical/computational research release. The package contains exact finite-dimensional theorems under stated models, executable verification, numerical proof-of-mechanism benchmarks, a prior-art… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/EVE-Synthetic-Dimension-Gauge-Field-Matter-Compiler.stockmatch-synthetic
StockMatch Synthetic Dataset & EDA Overview
Dataset Overview
This dataset contains 12,500 synthetic stock profiles designed for building an AI-powered stock recommendation system. It combines key numerical financial metrics with LLM-generated business summaries (Qwen2.5).
Exploratory Data Analysis (EDA) & Cleaning Summary
During the EDA process, we systematically analyzed and refined the dataset based on the following findings:
Irrelevant Columns… See the full description on the dataset page: https://huggingface.co/datasets/Kogann/stockmatch-synthetic.africa-synth-cerebral-palsy-synthetic-dataset-all
African Cerebral Palsy Synthetic Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-cerebral-palsy-synthetic-dataset-all.Aperture_Lab_Synthetic_Aperture_Sonar_v1
ApertureLab Synthetic SAS Dataset
Version 1.0 (September 2026). Author: Isaac Gerg. Made with
ApertureLab; samples, statistics and the
generation pipeline are described on the
dataset page.
1000 simulated synthetic aperture sonar (SAS) images, each an 80 m along-track
by 200 m range swath from a HISAS 1030-class 100 kHz sonar on a straight
track, beamformed by time-domain back-projection at 2.5 cm pixels and
delivered as dynamic-range-compressed (DRC) TIFF LZW images with COCO… See the full description on the dataset page: https://huggingface.co/datasets/idg101/Aperture_Lab_Synthetic_Aperture_Sonar_v1.africa-synth-concrete-compressive-strength-comoros
Africa Synth Concrete Compressive Strength Comoros | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: infrastructure_transport - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-concrete-compressive-strength-comoros.africa-synth-energy-oilgas-gas-production-nigeria
Africa Synth Energy Oilgas Gas Production Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-oilgas-gas-production-nigeria.synthetic-cpp
Dataset Card for Synthetic C++ Dataset
Dataset Description
Dataset Card for Synthetic C++ Dataset
Dataset Description
Homepage: [---
Dataset Card for Synthetic C++ Dataset
Dataset Description
Homepage: [https://huggingface.co/datasets/ReySajju742/synthetic-cpp/]
Point of Contact: [ReySajju742]
Dataset Summary
This dataset contains 10,000 rows of synthetically generated data focusing on the topic of "C++… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/synthetic-cpp.africa-synth-energy-oilgas-pipelines-nigeria
Africa Synth Energy Oilgas Pipelines Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-oilgas-pipelines-nigeria.Synthetic-ESCO-skill-sentences
Synthetic job ads for all ESCO skills
Dataset Summary
This dataset contains 10 synthetically generated job ad sentences for almost all (99.5%) skills in ESCO v1.1.0.
Languages
We use the English version of ESCO, and all generated sentences are in English.
Dataset Structure
The dataset consists of 138,260 (sentence, skill) pairs.
Citation Information
If you use this dataset, please include the following reference:… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Synthetic-ESCO-skill-sentences.smishing-syntheticIT-helpdesk-synthetic-ticketsafrica-synth-textile-garment-production-all
African Textile & Garment Production | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: csv - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-textile-garment-production-all.synthetic-indian-logical-reasoning-CoTyes
tube-cad-synthetic
Synthetic Tube CAD
22,800 generated CAD parts. Four classes. Original geometry and labels.
Identify round, square and rectangular tubes, plus confusing non-tubes. The parts include bends, holes, slots and cut ends.
Browse the files · W&B experiments · The story behind the dataset
Choose a collection
Main collection
Targeted collection
Start here if…
You want the broad original dataset
You want additional difficult variations
Parts
20,000
2,800… See the full description on the dataset page: https://huggingface.co/datasets/Ayt1da/tube-cad-synthetic.serena-synthetic-it-27h
Qwen3-TTS Italian Synthetic Speech (27h)
Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with
Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV,
Piper-ready metadata.
Dataset summary
Property
Value
Clips (train / val)
26,523 / 2,947
Total duration
~27.3 h (98,099 s)
Sample rate
22,050 Hz mono, 16-bit WAV
Loudness
Normalized to -23 LUFS, silence-trimmed
Language
Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-27h.synthetic_multilingual_llm_prompts
Image generated by DALL-E. See prompt for more details
📝🌐 Synthetic Multilingual LLM Prompts
Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.synthetic-intent-classifier-dataset-v1
Tanaos Intent Classifier Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate intent classification systems — models that can identify user intents in conversational AI applications, such as chatbots and virtual assistants.
Our flagship intent classification model, tanaos-intent-classifier-v1, was trained on this dataset.
Dataset Summary
The dataset contains text… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-intent-classifier-dataset-v1.reconriver-synthetic-reconciliation
ReconRiver synthetic reconciliation dataset
This dataset contains deterministic, entirely synthetic payment-reconciliation scenarios generated by the project's Java 21 dataset generator. Each scenario starts with a canonical payment collection and derives an internal ledger, processor events, bank settlements, and expected reconciliation ground truth.
No record describes a genuine person, payment, card, bank account, address, or customer. Identifiers use a conspicuous SYNTH-… See the full description on the dataset page: https://huggingface.co/datasets/heybadrinath/reconriver-synthetic-reconciliation.Crop-Disease-Image-Eval-Synthetic
Crop, Category, Disease and Pest Test Set
11,057 smallholder-farmer photographs sent to FarmerChat from Ethiopia, India, Kenya and Nigeria, each
labelled with the crop, whether the problem is a disease or a pest, and which one. This is the held-out
test split of a four-head classification benchmark, restricted to the rows whose labels came from an
independent model council rather than from the production vendor.
Why 11,057 and not 16,275
The full held-out split is… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/Crop-Disease-Image-Eval-Synthetic.africa-synth-autism-dataset-all
African Autism Spectrum Disorder Synthetic Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-autism-dataset-all.synthetic-legal
⚖️ Synthetic Legal (Query, Response) Dataset
📚 140,000 synthetic (legal query, legal response) pairs across 13 legal domains, built to resemble the structure of real-world fact patterns and citation-backed answers.
⚠️ Disclaimer: All text is synthetically generated and IS NOT LEGALLY ACCURATE. Citations are real but assigned at random, and the verified_solution and verification_method columns are template labels, not evidence of review. This dataset is not legal advice.… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/synthetic-legal.
