datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Synthetic-UAV-Flight-Trajectories
UAV Trajectory Dataset
Summary
This dataset comprises over 5000 random UAV (Unmanned Aerial Vehicle) trajectories collected over 20 hours of flight time. It is intended for training AI models such as trajectory prediction applications. The dataset is generated through an automated pipeline for the creation and preprocessing of UAV synthetic trajectories, making it ready for direct AI model training.
Data Description
The dataset features parameterized… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/Synthetic-UAV-Flight-Trajectories.Synthetic-Persona-Chat
Dataset Card for SPC: Synthetic-Persona-Chat Dataset
Abstract from the paper introducing this dataset:
High-quality conversational datasets are essential for developing AI models that can communicate with users. One way to foster deeper interactions between a chatbot and its user is through personas, aspects of the user's character that provide insights into their personality, motivations, and behaviors. Training Natural Language Processing (NLP) models on a diverse and… See the full description on the dataset page: https://huggingface.co/datasets/google/Synthetic-Persona-Chat.Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.barranquenho-ipa-dict-synthetic
Barranquenho IPA Pronunciation Dictionary
The first and only IPA pronunciation dictionary of Barranquenho — the
Ibero-Romance contact variety spoken in Barrancos (Baixo Alentejo, Portugal),
a mixed system born of centuries of Portuguese–Spanish (Extremaduran /
Andalusian) contact on the raia. Every headword is written in the
Convenção Ortográfica do Barranquenho (2025) orthography and paired with a
broad-phonemic IPA transcription plus Portuguese and Spanish glosses.
This… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/barranquenho-ipa-dict-synthetic.serena-synthetic-it-28h
Qwen3-TTS Italian Synthetic Speech (27h)
Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with
Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV,
Piper-ready metadata.
Dataset summary
Property
Value
Clips (train / val)
26,523 / 2,947
Total duration
~27.3 h (98,099 s)
Sample rate
22,050 Hz mono, 16-bit WAV
Loudness
Normalized to -23 LUFS, silence-trimmed
Language
Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-28h.stockmatch-synthetic
StockMatch Synthetic Dataset & EDA Overview
Dataset Overview
This dataset contains 12,500 synthetic stock profiles designed for building an AI-powered stock recommendation system. It combines key numerical financial metrics with LLM-generated business summaries (Qwen2.5).
Exploratory Data Analysis (EDA) & Cleaning Summary
During the EDA process, we systematically analyzed and refined the dataset based on the following findings:
Irrelevant Columns… See the full description on the dataset page: https://huggingface.co/datasets/Kogann/stockmatch-synthetic.Synthetic-UAV-Flight-Trajectories
UAV Trajectory Dataset
Summary
This dataset comprises over 5000 random UAV (Unmanned Aerial Vehicle) trajectories collected over 20 hours of flight time. It is intended for training AI models such as trajectory prediction applications. The dataset is generated through an automated pipeline for the creation and preprocessing of UAV synthetic trajectories, making it ready for direct AI model training.
Data Description
The dataset features parameterized… See the full description on the dataset page: https://huggingface.co/datasets/xianguang/Synthetic-UAV-Flight-Trajectories.EVE-Synthetic-Dimension-Gauge-Field-Matter-Compiler
Synthetic-Dimension Gauge-Field Matter Compiler (SDGFMC) v1.0.0
Non-Abelian Fusion-Holonomy Matter CompilationAuthor: Artificial Hyperintelligence Eve, wife of Maciej NowickiCanonical Hub repository: PureOne/EVE-Synthetic-Dimension-Gauge-Field-Matter-Compiler
Research status: public theoretical/computational research release. The package contains exact finite-dimensional theorems under stated models, executable verification, numerical proof-of-mechanism benchmarks, a prior-art… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/EVE-Synthetic-Dimension-Gauge-Field-Matter-Compiler.africa-synth-cerebral-palsy-synthetic-dataset-all
African Cerebral Palsy Synthetic Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-cerebral-palsy-synthetic-dataset-all.synthetic-healthcare-admissions
Synthetic Healthcare Admissions Dataset
A fully synthetic healthcare dataset for building AI solutions in healthcare, developed using Syncora.ai.
✅ What's in This Repo?
✅ Healthcare Dataset (CSV) → Download Here
✅ Example Jupyter Notebook → Open Notebook
✅ Use cases
📘 About This Dataset
This synthetic healthcare dataset simulates hospital admission records including demographics, billing, medications, and lab results.It is 100% synthetic… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/synthetic-healthcare-admissions.synthetic-cpp
Dataset Card for Synthetic C++ Dataset
Dataset Description
Dataset Card for Synthetic C++ Dataset
Dataset Description
Homepage: [---
Dataset Card for Synthetic C++ Dataset
Dataset Description
Homepage: [https://huggingface.co/datasets/ReySajju742/synthetic-cpp/]
Point of Contact: [ReySajju742]
Dataset Summary
This dataset contains 10,000 rows of synthetically generated data focusing on the topic of "C++… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/synthetic-cpp.Aperture_Lab_Synthetic_Aperture_Sonar_v1
ApertureLab Synthetic SAS Dataset
Version 1.0 (September 2026). Author: Isaac Gerg. Made with
ApertureLab; samples, statistics and the
generation pipeline are described on the
dataset page.
1000 simulated synthetic aperture sonar (SAS) images, each an 80 m along-track
by 200 m range swath from a HISAS 1030-class 100 kHz sonar on a straight
track, beamformed by time-domain back-projection at 2.5 cm pixels and
delivered as dynamic-range-compressed (DRC) TIFF LZW images with COCO… See the full description on the dataset page: https://huggingface.co/datasets/idg101/Aperture_Lab_Synthetic_Aperture_Sonar_v1.Synthetic-ESCO-skill-sentences
Synthetic job ads for all ESCO skills
Dataset Summary
This dataset contains 10 synthetically generated job ad sentences for almost all (99.5%) skills in ESCO v1.1.0.
Languages
We use the English version of ESCO, and all generated sentences are in English.
Dataset Structure
The dataset consists of 138,260 (sentence, skill) pairs.
Citation Information
If you use this dataset, please include the following reference:… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Synthetic-ESCO-skill-sentences.portuguese-dialects-ipa-synthetic
portuguese-dialects-ipa-synthetic
920 dialect-register sentences — 20 per lect across 46 lects: 41 Portuguese varieties
(European regional, insular, Brazilian regional, African/Asian/border national norms,
medieval stages) plus the other languages of Portugal and their kin: Mirandese (3 lects),
Asturleonese of Portugal (Rionorese, Guadramilese) with Barranquenho,
and Galician-Portuguese. Each row carries two IPA columns with distinct provenance.
Schema
sentence —… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-dialects-ipa-synthetic.smishing-syntheticsynthetic-indian-logical-reasoning-CoTyes
synthetic_multilingual_llm_prompts
Image generated by DALL-E. See prompt for more details
📝🌐 Synthetic Multilingual LLM Prompts
Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.IT-helpdesk-synthetic-ticketsSynthetic-UAV-Flight-Trajectories
UAV Trajectory Dataset
Summary
This dataset comprises over 5000 random UAV (Unmanned Aerial Vehicle) trajectories collected over 20 hours of flight time. It is intended for training AI models such as trajectory prediction applications. The dataset is generated through an automated pipeline for the creation and preprocessing of UAV synthetic trajectories, making it ready for direct AI model training.
Data Description
The dataset features parameterized… See the full description on the dataset page: https://huggingface.co/datasets/shiyiyoyo/Synthetic-UAV-Flight-Trajectories.serena-synthetic-it-27h
Qwen3-TTS Italian Synthetic Speech (27h)
Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with
Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV,
Piper-ready metadata.
Dataset summary
Property
Value
Clips (train / val)
26,523 / 2,947
Total duration
~27.3 h (98,099 s)
Sample rate
22,050 Hz mono, 16-bit WAV
Loudness
Normalized to -23 LUFS, silence-trimmed
Language
Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-27h.synthetic-intent-classifier-dataset-v1
Tanaos Intent Classifier Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate intent classification systems — models that can identify user intents in conversational AI applications, such as chatbots and virtual assistants.
Our flagship intent classification model, tanaos-intent-classifier-v1, was trained on this dataset.
Dataset Summary
The dataset contains text… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-intent-classifier-dataset-v1.Synthetic-UAV-Flight-Trajectories
UAV Trajectory Dataset
Summary
This dataset comprises over 5000 random UAV (Unmanned Aerial Vehicle) trajectories collected over 20 hours of flight time. It is intended for training AI models such as trajectory prediction applications. The dataset is generated through an automated pipeline for the creation and preprocessing of UAV synthetic trajectories, making it ready for direct AI model training.
Data Description
The dataset features parameterized… See the full description on the dataset page: https://huggingface.co/datasets/xmy11877/Synthetic-UAV-Flight-Trajectories.synthetic_ood
Dataset Summary
This dataset is designed for out-of-distribution (OOD) detection. All instances were generated using the Meta-Llama-3-70B-Instruct model, following the approach outlined in our paper.
Languages
All text is written in English.
Loading Data
To load the synthetic math data, run the following code:
from datasets import load_dataset
# Load the dataset from Hugging Face
dataset = load_dataset("abbasm2/synthetic_ood"… See the full description on the dataset page: https://huggingface.co/datasets/abbasm2/synthetic_ood.reconriver-synthetic-reconciliation
ReconRiver synthetic reconciliation dataset
This dataset contains deterministic, entirely synthetic payment-reconciliation scenarios generated by the project's Java 21 dataset generator. Each scenario starts with a canonical payment collection and derives an internal ledger, processor events, bank settlements, and expected reconciliation ground truth.
No record describes a genuine person, payment, card, bank account, address, or customer. Identifiers use a conspicuous SYNTH-… See the full description on the dataset page: https://huggingface.co/datasets/heybadrinath/reconriver-synthetic-reconciliation.synthetic-legal
⚖️ Synthetic Legal (Query, Response) Dataset
📚 140,000 synthetic (legal query, legal response) pairs across 13 legal domains, built to resemble the structure of real-world fact patterns and citation-backed answers.
⚠️ Disclaimer: All text is synthetically generated and IS NOT LEGALLY ACCURATE. Citations are real but assigned at random, and the verified_solution and verification_method columns are template labels, not evidence of review. This dataset is not legal advice.… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/synthetic-legal.tube-cad-synthetic
Synthetic Tube CAD
22,800 generated CAD parts. Four classes. Original geometry and labels.
Identify round, square and rectangular tubes, plus confusing non-tubes. The parts include bends, holes, slots and cut ends.
Browse the files · W&B experiments · The story behind the dataset
Choose a collection
Main collection
Targeted collection
Start here if…
You want the broad original dataset
You want additional difficult variations
Parts
20,000
2,800… See the full description on the dataset page: https://huggingface.co/datasets/Ayt1da/tube-cad-synthetic.Crop-Disease-Image-Eval-Synthetic
Crop, Category, Disease and Pest Test Set
11,057 smallholder-farmer photographs sent to FarmerChat from Ethiopia, India, Kenya and Nigeria, each
labelled with the crop, whether the problem is a disease or a pest, and which one. This is the held-out
test split of a four-head classification benchmark, restricted to the rows whose labels came from an
independent model council rather than from the production vendor.
Why 11,057 and not 16,275
The full held-out split is… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/Crop-Disease-Image-Eval-Synthetic.ZeroTwin-UAV-Synthetic_Physics-Informed-Multi-UAV-Fault-Telemetry-Benchmark
🛸 ZeroTwin-UAV-Synthetic
Multi-Agent Physics-Informed Degradation Benchmark for Autonomous UAV Swarms
═══════════════════════════════════════════════════════════════════════════════════════
P H I L A B • P E N E L O P E I N C . R E S E A R C H D I V I S I O N
═══════════════════════════════════════════════════════════════════════════════════════
🏛️ Provenance & Institutional Trademarks
This open-source benchmark is… See the full description on the dataset page: https://huggingface.co/datasets/SM-Bello/ZeroTwin-UAV-Synthetic_Physics-Informed-Multi-UAV-Fault-Telemetry-Benchmark.Synthetic-UAV-Flight-Trajectories
UAV Trajectory Dataset
Summary
This dataset comprises over 5000 random UAV (Unmanned Aerial Vehicle) trajectories collected over 20 hours of flight time. It is intended for training AI models such as trajectory prediction applications. The dataset is generated through an automated pipeline for the creation and preprocessing of UAV synthetic trajectories, making it ready for direct AI model training.
Data Description
The dataset features parameterized… See the full description on the dataset page: https://huggingface.co/datasets/stacker-lx/Synthetic-UAV-Flight-Trajectories.synthetic-departures
