datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
enron-emailsenron-qa-emails-dasovich-jguertin-mcro-forensic-corpus-emails
Guertin MCRO Forensic Corpus: Emails
Contents: 352 email messages (.eml) in 11 categories — LinkedIn search-appearance notifications (99); messages before 2023-01-21 (82); messages after 2023-01-21 (76); correspondence with the first Rule 20 examiner (26); correspondence with the public defender (41); delivery of the March 5, 2025 hearing transcript (1); the Minnesota Attorney General's office (federal case) (1); 2026 correspondence (15); U.S. Senator Amy Klobuchar's office… See the full description on the dataset page: https://huggingface.co/datasets/Matt1up/guertin-mcro-forensic-corpus-emails.enron_aeslc_emailshillary-clinton-emails-wikileaksenron_emails_sample_questionsMarketing-Emails
Marketing Emails
A curated corpus of synthetically generated yet realistic marketing email messages designed to support research in Domain Adaptation, Natural Language Processing (NLP), Data Science, Machine Learning, and Communication research.
The dataset is appropriate for a wide spectrum of training paradigms—including pre-training, fine-tuning, and domain adaptation—as well as for rigorous evaluation of models targeting domain-specific language understanding and generation… See the full description on the dataset page: https://huggingface.co/datasets/marketeam/Marketing-Emails.enron_emails_parsedenrun-emails-token-classificationemail-spam-classification
Email Spam Classification
The dataset consists of a collection of emails categorized into two major classes: spam and not spam. It is designed to facilitate the development and evaluation of spam detection or email filtering systems.
The spam emails in the dataset are typically unsolicited and unwanted messages that aim to promote products or services, spread malware, or deceive recipients for various malicious purposes. These emails often contain misleading subject lines… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/email-spam-classification.marketing_emailsenrun-emails-text-classificationepstein-emails
Epstein Email Messages Dataset
Dataset Summary
This dataset contains 4,272 individual email messages extracted directly from screenshot images using advanced Vision LLM technology. Unlike other datasets that work with only pre-OCR'd text, this dataset also used the original JPG screenshot images from the U.S. House Oversight Committee release using Qwen 2.5 VL 72B vision model to extract structured email data with high accuracy.
The dataset powers a live Progressive Web… See the full description on the dataset page: https://huggingface.co/datasets/to-be/epstein-emails.epstein-emails
Epstein Email Threads Dataset
Dataset Summary
This dataset contains 5,082 parsed email threads extracted from OCR'd documents released by the U.S. House Oversight Committee. The emails have been processed using large language models to extract structured information including senders, recipients, timestamps, subjects, and message bodies, with OCR errors corrected and footers removed.
Dataset Description
Overview
This is a structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/notesbymuneeb/epstein-emails.email_spam-qwen3-vl-32benron-emailsro-business-emailsPhishing_emails_testphishing_emails-data
🛡️ Phishing Email Classification Dataset
This dataset is curated for fine-tuning LLMs on the task of phishing email detection. It originates from this Kaggle dataset and has been transformed to better suit LLM-based classification tasks.
📦 Dataset Features
Each row is a labeled email, with either:
safe email (label = 0)
phishing email (label = 1)
The dataset includes metadata (sender, receiver, date, subject) and cleaned email body.
Two main columns:
Email Text:… See the full description on the dataset page: https://huggingface.co/datasets/drorrabin/phishing_emails-data.Panza-emails
The Panza Emails dataset
This dataset contains collections of emails of three authentic users (david, isabel, and marcus), with personal information (names, places, etc.) replaced by other ones for donor privacy.
Except for these changes, the language of the emails is genuine. The intention of this dataset is to allow researchers to study strategies for text personalization.
The data was donated explicitly for this purpose. This dataset is ethically collected and fully licensed for… See the full description on the dataset page: https://huggingface.co/datasets/ISTA-DASLab/Panza-emails.email-spam-filterepstein-emails
Epstein Email Threads Dataset
Dataset Summary
This dataset contains 5,082 parsed email threads extracted from OCR'd documents released by the U.S. House Oversight Committee. The emails have been processed using large language models to extract structured information including senders, recipients, timestamps, subjects, and message bodies, with OCR errors corrected and footers removed.
Dataset Description
Overview
This is a structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/Hannah2704/epstein-emails.details_postbot__pythia-160m-hq-emails
Dataset Card for Evaluation run of postbot/pythia-160m-hq-emails
Dataset Summary
Dataset automatically created during the evaluation run of model postbot/pythia-160m-hq-emails on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_postbot__pythia-160m-hq-emails.spam-ham-phish-emails-latestexcel-huggingface-emails-1680-product-catalog
Customer Satisfaction Survey 2024
Customer satisfaction survey responses collected from the mobile app between January and December 2024.
Columns
response_id
customer_id
satisfaction_score
churn_risk
survey_date
scaling_mia_the_pile_00_Enron_Emailspile-enron_emails
Dataset Creation Process
These subsets were created by streaming over the rows from monology/pile-uncopyrighted and filtering by the meta column. Each subset is generally limited to the first 100,000 qualifying rows encountered.
Citations
If you use this dataset, please cite the original Pile papers:
@article{gao2020pile,
title={The Pile: An 800GB dataset of diverse text for language modeling},
author={Gao, Leo and Biderman, Stella and Black, Sid and Golding, Laurence and… See the full description on the dataset page: https://huggingface.co/datasets/timaeus/pile-enron_emails.email-security
EMAIL_SECURITY
A preference dataset for EMAIL_SECURITY, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate.
Format
Standard preference / DPO schema — each row:
column
meaning
prompt
the request (originally prompt)
chosen
the human-preferred response
rejected
a worse response to the same prompt
source
the dataset/URL the row was harvested from
Splits
80/10/10 train… See the full description on the dataset page: https://huggingface.co/datasets/316usman/email-security.phishing-emails-multilingual
Phishing Emails Multilingual (ID/EN) — Synthetic
Dataset sintetis & edukatif 600 email dwibahasa Indonesia 🇮🇩 & English 🇺🇸 untuk riset deteksi phishing — oleh Febriyansyah.
⚠️ Synthetic & edu-defense-only — dibuat untuk pembelajaran defensive security, bukan untuk kampanye nyata. Jangan gunakan untuk aktivitas ilegal.
Ringkasan
600 baris — 300 phishing / 300 benign (seimbang), 321 EN / 279 ID
Kolom: id (int), language (id/en), text (string, badan email)… See the full description on the dataset page: https://huggingface.co/datasets/Febriyansyah/phishing-emails-multilingual.email-safety-triage-10k
Email Safety Triage 10k
This dataset contains 10,000 supervised examples for classifying email and email-adjacent content for operational triage, phishing/spam risk, and prompt-attack filtering.
Each JSONL row has two string fields:
input: an instruction plus email, security-review text, or prompt/email fragment.
output: compact strict JSON with triage, priority, risk, should_process, confidence, and reason.
The dataset is intended for fine-tuning and evaluating classifiers… See the full description on the dataset page: https://huggingface.co/datasets/weijianzhg/email-safety-triage-10k.
