datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
africa-synth-cerebral-palsy-synthetic-dataset-all
African Cerebral Palsy Synthetic Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-cerebral-palsy-synthetic-dataset-all.synthetic-intent-classifier-dataset-v1
Tanaos Intent Classifier Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate intent classification systems — models that can identify user intents in conversational AI applications, such as chatbots and virtual assistants.
Our flagship intent classification model, tanaos-intent-classifier-v1, was trained on this dataset.
Dataset Summary
The dataset contains text… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-intent-classifier-dataset-v1.computer-science-synthetic-datasetos-rfodg-outdoor-uav-synthetic-dataset-taif-saudi-arabia
UAV Trajectory Simulation Dataset for Terrain-Based Localization
Dataset Overview
This dataset contains simulated UAV flight data generated using ROS2, Gazebo, and PX4 autopilot system. The dataset features a quadcopter performing autonomous flight trajectories over realistic terrain imported from satellite imagery and Digital Elevation Model (DEM) maps of the Taif region in Saudi Arabia.
Dataset Files
The dataset contains:
7 trajectory CSV files:… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/os-rfodg-outdoor-uav-synthetic-dataset-taif-saudi-arabia.synthetic-financial-data
Dataset Card for Dataset Name
Extracted from Kaggle: https://www.kaggle.com/datasets/ealaxi/paysim1?resource=download
Purpose is to have synthetic data on here for testing.
Synthetic-Diabetes-Dataset
⛽ Synthetic Diabetes Data
This dataset contains various features that are helpful for predicting if a patient has diabetes. The data is compiled into a single csv file for analysis and model training.
📁 Dataset Description
Contains synthetic data on based on synthetic patient information.
Columns
The dataset includes the following columns:
abdominal_obesity (int)
alcohol_consumption_per_week (int):
The amount of alcohol consumption per week of the patient.… See the full description on the dataset page: https://huggingface.co/datasets/MaxPrestige/Synthetic-Diabetes-Dataset.synthetic-emotion-detection-dataset-v1
Tanaos Emotion Detection Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate emotion detection systems — models that classify the main emotion expressed in text as one of eight possible categories: joy, anger, fear, sadness, surprise, disgust, excitement, or neutral. It can be used to build emotion detection models for various applications, such as customer feedback analysis, social… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-emotion-detection-dataset-v1.synthetic-patients-FHIR-data
Synthetic patient FHIR data
Synthetic patient records generated with Synthea
and exported from a HAPI FHIR store.
No real patient data. Every record here is machine generated. Nothing in this
dataset derives from a real person.
Patients: 935
Cohorts: 11
FHIR resources: 1,004,084 (FHIR R4)
The same patients are provided in two shapes. Pick whichever fits your work.
The dataset viewer above shows the flat table, which is the row-oriented view.
The bundles are whole FHIR… See the full description on the dataset page: https://huggingface.co/datasets/anas-elghafari/synthetic-patients-FHIR-data.NL-to-LTL-Synthetic-Datasetsynthetic-spam-detection-dataset-german
Tanaos Spam Detection German Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate spam detection systems — models that detect, classify, or filter unsolicited commercial advertisement, fraudulent messages, or other unwanted content in text form — in German.
Our german spam detection model, tanaos-spam-detection-german, was trained on this dataset.
Dataset Summary
The… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-spam-detection-dataset-german.Synthetic_dataset-MCAeConsultation_Sentiment_AnalysisBangladesh-Voter-Synthetic-Dataset
🗳️ Bangladesh Voter Dataset
📜 Dataset Description
The Bangladesh Voter Dataset is a synthetic dataset containing voter information for the purpose of demonstrating data generation and processing techniques. Each voter record includes both Bengali and English names, gender, NID, address, and profile information.
📊 Dataset Structure
The dataset is structured as follows:
profile: A URL to the voter's profile image.
nid: A unique National Identification Number.… See the full description on the dataset page: https://huggingface.co/datasets/jonybepary/Bangladesh-Voter-Synthetic-Dataset.guide-to-level-measurement-2021-edition-en-77708_combined_synthetic_datasynthetic-NER-dataset-v1
Tanaos NER Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate Named Entity Recognition (NER) systems — models that identify and classify named entities in text into predefined categories such as PERSON, ORG, LOCATION, DATE, and more. It can be used to train NER models from scratch or fine-tune existing ones.
Our flagship NER model, tanaos-NER-v1, was trained on this dataset.… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-NER-dataset-v1.uk_retail_store_synthetic_dataset
Synthetic Data Generation Demo — UK Retail Dataset
Welcome to this synthetic data generation demo repository by Syncora.ai. This project showcases how to generate synthetic data using real-world tabular structures, demonstrated on a UK retail dataset with columns such as:
Country
CustomerID
UnitPrice
InvoiceDate
Quantity
StockCode
This dataset is designed for dataset for LLM training and AI development, enabling developers to work with privacy-safe, high-quality… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/uk_retail_store_synthetic_dataset.synthetic-guardrail-dataset-v1
Tanaos Guardrail Training Dataset
[!CAUTION]
We now have a newer version of this dataset: tanaos/synthetic-guardrail-dataset-v2 with improved coverage and quality. Consider using that instead.
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate guardrail systems — models that detect, classify, or filter unsafe, harmful, or policy-violating text content. It can be used to train moderation models or… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-guardrail-dataset-v1.translated-dataset-synthetic-retrieval-taskssynthetic-topic-classification-dataset-v1
Tanaos Topic Classification Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate Topic Classification models — models that can classify text into predefined topics or categories.
Our flagship Topic Classification model, tanaos-topic-classification-v1, was trained on this dataset.
Dataset Summary
The dataset contains text samples labeled with their corresponding topics.… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-topic-classification-dataset-v1.synthetic-topic-classification-dataset-v1
Tanaos Topic Classification Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate Topic Classification models — models that can classify text into predefined topics or categories.
Our flagship Topic Classification model, tanaos-topic-classification-v1, was trained on this dataset.
Dataset Summary
The dataset contains text samples labeled with their corresponding topics.… See the full description on the dataset page: https://huggingface.co/datasets/donajui/synthetic-topic-classification-dataset-v1.player-valuer-synthetic-datasynthetic-sentiment-analysis-dataset-v1
Tanaos Sentiment Analysis Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate sentiment analysis systems — models that classify the sentiment expressed in text as one of five possible categories: very_negative, negative, neutral, positive or very_positive. It can be used to build sentiment analysis models for various applications, such as customer feedback analysis, social media… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-sentiment-analysis-dataset-v1.synthetic-spam-detection-dataset-spanish
Tanaos Spam Detection Spanish Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate spam detection systems — models that detect, classify, or filter unsolicited commercial advertisement, fraudulent messages, or other unwanted content in text form — in Spanish.
Our spanish spam detection model, tanaos-spam-detection-spanish, was trained on this dataset.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-spam-detection-dataset-spanish.synthetic-medical-ehr-dataset
🏥 Synthetic Privacy-Preserving Medical EHR Dataset
Dataset Description
A fully synthetic collection of 10,000 Electronic Health Records (EHRs) for binary classification research. The task is predicting adverse patient outcomes — deterioration or death — from clinical, demographic, and vital sign features. No real patient data was used at any stage. Privacy-safe under GDPR, CCPA, and HIPAA.
Supported Tasks
tabular-classification: Predict adverse_outcome (0/1)… See the full description on the dataset page: https://huggingface.co/datasets/mdsajjadullah/synthetic-medical-ehr-dataset.synthetic-guardrail-dataset-v2
Tanaos Guardrail Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate guardrail systems — models that detect, classify, or filter unsafe, harmful or potentially dangerous content. It can be used to train moderation models or integrate LLM safety filters for applications like chatbots, content generation, and user-facing AI systems.
Our flagship guardrail model, tanaos-guardrail-v2… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-guardrail-dataset-v2.Law_domain_synthetic_dataSynthetic-Diabetes-Dataset
⛽ Synthetic Diabetes Data
This dataset contains various features that are helpful for predicting if a patient has diabetes. The data is compiled into a single csv file for analysis and model training.
📁 Dataset Description
Contains synthetic data on based on synthetic patient information.
Columns
The dataset includes the following columns:
abdominal_obesity (int)
alcohol_consumption_per_week (int):
The amount of alcohol consumption per week of the patient.… See the full description on the dataset page: https://huggingface.co/datasets/dionysusss/Synthetic-Diabetes-Dataset.synthetic-cameroon-national-id-card-orc-dataset
🇨🇲 Cameroon National ID Card OCR Dataset
Synthetic dataset for information extraction from Cameroonian National Identity Cards via OCR.
📋 Description
This dataset contains 60,000 examples of simulated OCR text from Cameroonian National ID Cards with corresponding structured information in JSON format. The data covers two ID card formats (2018 and 2025) and two sides (front/back) with different levels of OCR noise.
🎯 Use Cases
Fine-tuning LLM models for… See the full description on the dataset page: https://huggingface.co/datasets/thekfp/synthetic-cameroon-national-id-card-orc-dataset.It-support-synthetic-dataJailbreaking-Synthetic-Datasetswarm-up_synthetic-data
