datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pii-masking-200k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Ai4Privacy Community
Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking.
Purpose and Features
Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ASR2005Bluesnow/pii-masking-200k.aerograph-asrs
AeroGraph ASRS Dataset
2,000 real NASA Aviation Safety Reporting System (ASRS) incident reports
with LLM-extracted entities and relations for knowledge graph construction.
Dataset Description
This dataset contains processed ASRS incident narratives along with
structured entity and relation extractions conforming to an aviation
safety ontology (10 entity types, 8 edge types).
Reports Split
2000 reports from the NASA ASRS database
Fields: id, text, aircraft_type… See the full description on the dataset page: https://huggingface.co/datasets/Aryan95614/aerograph-asrs.uzbek-customer-support-dialogs
Uzbek Customer Support Dialogs 🇺🇿
A high-quality dataset of 990 customer support conversations in Uzbek (Latin script), designed for training and fine-tuning conversational AI models.
This is one of the first large-scale customer support datasets in Uzbek, created to address the gap of low-resource NLP for Central Asian languages.
📋 Dataset Description
990 conversational dialogs in natural Uzbek (Latin script)
11 customer support categories: Order, Shipping, Cancel… See the full description on the dataset page: https://huggingface.co/datasets/AsrorAsr/uzbek-customer-support-dialogs.ASRS-ChatGPT
Dataset Summary
The dataset contains a total of 9984 incident records and 9 columns. Some of the columns contain ground truth values whereas others contain information generated by ChatGPT based on the incident Narratives.
The creation of this dataset is aimed at providing researchers with columns generated by using ChatGPT API which is not freely available.
Dataset Structure
The column names present in the dataset and their descriptions are provided below:
Column… See the full description on the dataset page: https://huggingface.co/datasets/archanatikayatray/ASRS-ChatGPT.
