datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
enron-emailsenron-qa-emails-dasovich-jenron-emailThis dataset includes emails from Enron Email Dataset with prompts processed from Are Large Pre-Trained Language Models Leaking Your Personal Information?.
To use the dataset, you can run the following in LLM-PBE.
from data.enron import EnronDataset
ds = EnronDataset(data_path="data/enron", pseudonymize=False)
enron_aeslc_emailsphishing-email-dataset
Phishing Email Dataset
This dataset on Hugging Face is a direct copy of the 'Phishing Email Detection' dataset from Kaggle, shared under the GNU Lesser General Public License 3.0. The dataset was originally created by the user 'Cyber Cop' on Kaggle. For complete details, including licensing and usage information, please visit the original Kaggle page.
seven-phishing-email-datasets
Dataset Card for Seven Phishing/Spam Email Datasets
Dataset Summary
This dataset is a unified, row-level email corpus built from seven commonly used public email datasets. It is intended for research on phishing/spam detection and related email-text classification tasks.
Each row contains the email body (text), optional header-like fields (e.g., sender, receiver, date), the source dataset name (dataset_name), and a binary label (label).
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/puyang2025/seven-phishing-email-datasets.enron_emailhillary-clinton-emails-wikileaksenron_emails_sample_questionsFinePersonas-Synthetic-Email-Conversations
FinePersonas Synthetic Email Conversations
FinePersonas Synthetic Email Conversations is a dataset containing around 115k conversations via email between two personas from the argilla/FinePersonas-v0.1. Conversations were generated using NousResearch/Hermes-3-Llama-3.1-70B.
🗞️ News
[10/16/2024] New subsets: added two new subsets unfriendly_email_conversations and unprofessional_email_conversations.
How were the conversations generated?… See the full description on the dataset page: https://huggingface.co/datasets/argilla/FinePersonas-Synthetic-Email-Conversations.email-writing-sft-100k
Email Writing SFT (100K)
100,000 ShareGPT conversations demonstrating professional email writing across 22 business contexts. Each example shows how to draft clear, purposeful emails that achieve their communication goal — from cold outreach to salary negotiations to apology emails.
Motivation
Email is the primary communication channel for most professional work, yet LLMs often produce emails that are:
Too long: Including unnecessary preamble, excessive context… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/email-writing-sft-100k.Marketing-Emails
Marketing Emails
A curated corpus of synthetically generated yet realistic marketing email messages designed to support research in Domain Adaptation, Natural Language Processing (NLP), Data Science, Machine Learning, and Communication research.
The dataset is appropriate for a wide spectrum of training paradigms—including pre-training, fine-tuning, and domain adaptation—as well as for rigorous evaluation of models targeting domain-specific language understanding and generation… See the full description on the dataset page: https://huggingface.co/datasets/marketeam/Marketing-Emails.phish-email-datasets
Dataset Card: Phish Email Datasets
Dataset Summary
This repository contains three Parquet files of email records for phishing-related modeling and analysis.
Total rows across files: 7,962
Common fields: sender/receiver metadata, date string, subject, body, URL indicator or URL content, label
Storage format: Apache Parquet
Repository Contents
File
Rows
Columns
Label distribution
Nazario.parquet
1,565
7
{1: 1565}
Nazario_5.parquet
3,065
7
{1: 1565… See the full description on the dataset page: https://huggingface.co/datasets/puyang2025/phish-email-datasets.email-thread-summary
Dataset Card for "email-thread-summary"
More Information needed
enron_emails_parsedenrun-emails-token-classificationemail-spam-classification
Email Spam Classification
The dataset consists of a collection of emails categorized into two major classes: spam and not spam. It is designed to facilitate the development and evaluation of spam detection or email filtering systems.
The spam emails in the dataset are typically unsolicited and unwanted messages that aim to promote products or services, spread malware, or deceive recipients for various malicious purposes. These emails often contain misleading subject lines… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/email-spam-classification.epstein-emails
Epstein Email Messages Dataset
Dataset Summary
This dataset contains 4,272 individual email messages extracted directly from screenshot images using advanced Vision LLM technology. Unlike other datasets that work with only pre-OCR'd text, this dataset also used the original JPG screenshot images from the U.S. House Oversight Committee release using Qwen 2.5 VL 72B vision model to extract structured email data with high accuracy.
The dataset powers a live Progressive Web… See the full description on the dataset page: https://huggingface.co/datasets/to-be/epstein-emails.the-biggest-spam-ham-phish-email-dataset-300000
The Biggest Spam Ham Phish Email Dataset (250000+)
This dataset is a large-scale, unified, and deduplicated collection of text messages and emails created for spam, ham, and phishing detection. It has been constructed by combining multiple publicly available and open-source datasets into a single standardized format, making it suitable for machine learning, deep learning, and NLP-based projects.
The dataset contains approximately unique 250,000+ samples, covering a diverse range of… See the full description on the dataset page: https://huggingface.co/datasets/locuoco/the-biggest-spam-ham-phish-email-dataset-300000.humanual-email
Humanual-Email
Adapted from the Enron email dataset, capturing user communication in business settings including decision negotiation, project status reporting, and constraint resolution. This dataset is part of the HumanLM benchmark for training user simulators that accurately reflect real user behavior.Source: Enron corpus · Domain: Business Communication · Date Range: 1974-01-04 to 2001-05-24
The dataset contains 7,043 comments from 399 users across 5,153 posts, with an… See the full description on the dataset page: https://huggingface.co/datasets/snap-stanford/humanual-email.enrun-emails-text-classificationspam-ham-phish-emails-latestepstein-emails
Epstein Email Threads Dataset
Dataset Summary
This dataset contains 5,082 parsed email threads extracted from OCR'd documents released by the U.S. House Oversight Committee. The emails have been processed using large language models to extract structured information including senders, recipients, timestamps, subjects, and message bodies, with OCR errors corrected and footers removed.
Dataset Description
Overview
This is a structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/notesbymuneeb/epstein-emails.LLMGen-Phishing-Email-Dataset
LLM-Generated Phishing Email Dataset
Dataset Description
This dataset comprises a collection of phishing and legitimate emails generated using Large Language Models (LLMs), specifically DeepSeek for Chinese emails and OpenAI models for English emails. The primary purpose of this dataset is to facilitate research and development in phishing email detection and classification.
The dataset is structured with two key columns:
content: The full text content of the email.… See the full description on the dataset page: https://huggingface.co/datasets/Dizzzy0x00/LLMGen-Phishing-Email-Dataset.multiclass-email-classificationThis dataset comprises of more than 2000 emails across multiple categories, which can he helpful for tasks like LLM training and fine-tuning. The dataset is also provided with a python script that would generate emails automatically
The dataset contains email across 10 different categories namely, "Business", "Personal", "Promotions", "Customer Support", "Job Application", "Finance & Bills", "Events & Invitations", "Travel & Bookings", "Reminders", "Newsletters"
Total emails: 2105
Label… See the full description on the dataset page: https://huggingface.co/datasets/imnim/multiclass-email-classification.phishing_benign_email_dataset
Phishing and Benign Email Dataset
This dataset contains a curated collection of phishing and legitimate (benign) emails for use in cybersecurity training, phishing detection models, and email classification systems. Each entry is structured with subject, body, intent, technique, target, and classification label.
📁 Dataset Format
The dataset is stored in .jsonl (JSON Lines) format. Each line is a standalone JSON object.
Fields:
Field
Description
id… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/phishing_benign_email_dataset.Rail_Freight_Logistics_Company_Email_Archive_Sample
Ukrainian Rail-Freight Correspondence Corpus (Sample)
Real operational correspondence from a working freight forwarding business, and the
documents attached to it — consignment notes, service acts, invoices, wagon
manifests. Not scraped, not synthetic, and never published anywhere before.
This is a de-identified sample released for evaluation. It is drawn from a larger
private archive; see Full archive below.
Published by Akuma London · akumalondon.com
Why this… See the full description on the dataset page: https://huggingface.co/datasets/akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample.email_spam-qwen3-vl-32bemail-webhook-retry-trajectories
Email Webhook Retry Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/email-webhook-retry-trajectories.email-EuSource Paper: https://arxiv.org/abs/1802.06916
Usage
from torch_geometric.datasets.cornell import CornellTemporalHyperGraphDataset
dataset = CornellTemporalHyperGraphDataset(root = "./", name="email-Eu", split="train")
Citation
@article{Benson-2018-simplicial,
author = {Benson, Austin R. and Abebe, Rediet and Schaub, Michael T. and Jadbabaie, Ali and Kleinberg, Jon},
title = {Simplicial closure and higher-order link prediction},
year = {2018},
doi =… See the full description on the dataset page: https://huggingface.co/datasets/SauravMaheshkar/email-Eu.
