datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
phishing-email-dataset
Phishing Email Dataset
This dataset on Hugging Face is a direct copy of the 'Phishing Email Detection' dataset from Kaggle, shared under the GNU Lesser General Public License 3.0. The dataset was originally created by the user 'Cyber Cop' on Kaggle. For complete details, including licensing and usage information, please visit the original Kaggle page.
Marketing-Emails
Marketing Emails
A curated corpus of synthetically generated yet realistic marketing email messages designed to support research in Domain Adaptation, Natural Language Processing (NLP), Data Science, Machine Learning, and Communication research.
The dataset is appropriate for a wide spectrum of training paradigms—including pre-training, fine-tuning, and domain adaptation—as well as for rigorous evaluation of models targeting domain-specific language understanding and generation… See the full description on the dataset page: https://huggingface.co/datasets/marketeam/Marketing-Emails.email-spam-classification
Email Spam Classification
The dataset consists of a collection of emails categorized into two major classes: spam and not spam. It is designed to facilitate the development and evaluation of spam detection or email filtering systems.
The spam emails in the dataset are typically unsolicited and unwanted messages that aim to promote products or services, spread malware, or deceive recipients for various malicious purposes. These emails often contain misleading subject lines… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/email-spam-classification.the-biggest-spam-ham-phish-email-dataset-300000
The Biggest Spam Ham Phish Email Dataset (250000+)
This dataset is a large-scale, unified, and deduplicated collection of text messages and emails created for spam, ham, and phishing detection. It has been constructed by combining multiple publicly available and open-source datasets into a single standardized format, making it suitable for machine learning, deep learning, and NLP-based projects.
The dataset contains approximately unique 250,000+ samples, covering a diverse range of… See the full description on the dataset page: https://huggingface.co/datasets/locuoco/the-biggest-spam-ham-phish-email-dataset-300000.LLMGen-Phishing-Email-Dataset
LLM-Generated Phishing Email Dataset
Dataset Description
This dataset comprises a collection of phishing and legitimate emails generated using Large Language Models (LLMs), specifically DeepSeek for Chinese emails and OpenAI models for English emails. The primary purpose of this dataset is to facilitate research and development in phishing email detection and classification.
The dataset is structured with two key columns:
content: The full text content of the email.… See the full description on the dataset page: https://huggingface.co/datasets/Dizzzy0x00/LLMGen-Phishing-Email-Dataset.phishing-email-balanced-6000
Balanced Phishing Email Detection Subset
This dataset is a derived, randomly sampled subset of Cyber Cop's
Phishing Email Detection
dataset on Kaggle. The original dataset is distributed under the GNU Lesser
General Public License 3.0.
Dataset structure
The file phishing_email_subset.csv contains 6,000 English email examples:
text: email text.
label: 0 for a safe email and 1 for a phishing email.
Label
Class
Examples
0
Safe email
3,000
1
Phishing… See the full description on the dataset page: https://huggingface.co/datasets/jhonrayo99/phishing-email-balanced-6000.email-politeness-corpus
Email Politeness Corpus
This dataset accompanies the paper:
A Synthetic Request–Reply Email Corpus Annotated with Document-Level Politeness and Sentence-Level Face Acts
Roshad Alipanah, Valentin Barriere, and Jorge BaierFindings of the Association for Computational Linguistics: EMNLP 2026
The corpus consists of synthetic request–reply emails jointly annotated at two levels:
Sentence level: multi-label Face Act annotations grounded in Brown and Levinson's politeness theory.… See the full description on the dataset page: https://huggingface.co/datasets/ralipanah/email-politeness-corpus.email-spam-filteremail-campaigns
Email Marketing Campaign Analytics Dataset (Free Sample)
This is a free sample with 4,025 rows. The full dataset has 56,462 rows across 4 tables.
Email campaign performance data for a simulated B2B SaaS company running
120 campaigns over 18 months. 15,000 subscribers across 5 segments,
40,000 email events (sends, opens, clicks, bounces, unsubscribes).
Features realistic engagement curves: declining open rates over time,
segment-specific behavior, A/B test results, and two… See the full description on the dataset page: https://huggingface.co/datasets/mindweave/email-campaigns.enron_labeled_emails_with_subjects-llama2-7b_finetuningphishing-emails-multilingual
Phishing Emails Multilingual (ID/EN) — Synthetic
Dataset sintetis & edukatif 600 email dwibahasa Indonesia 🇮🇩 & English 🇺🇸 untuk riset deteksi phishing — oleh Febriyansyah.
⚠️ Synthetic & edu-defense-only — dibuat untuk pembelajaran defensive security, bukan untuk kampanye nyata. Jangan gunakan untuk aktivitas ilegal.
Ringkasan
600 baris — 300 phishing / 300 benign (seimbang), 321 EN / 279 ID
Kolom: id (int), language (id/en), text (string, badan email)… See the full description on the dataset page: https://huggingface.co/datasets/Febriyansyah/phishing-emails-multilingual.email-authentication
DMARC and SPF Adoption Among Large Organizations
Overview
This dataset records which of 36,120 large organizations publish SPF and DMARC records on their primary domain, with firmographic context for each: industry, employee band, country, locality and founding year.
SPF lists the servers allowed to send mail for a domain. DMARC tells receiving servers what to do with mail that fails that check, and where to send reports. A domain with SPF but no DMARC has… See the full description on the dataset page: https://huggingface.co/datasets/stackscan/email-authentication.fake-email-campaignEmail_Intent_Classification
Dataset Information
This is a dataset of English sentences used in emails with six basic categories: request, informational, transaction, feedback.
An example looks as follows: {"Email": "Your subscription renewal is confirmed. Thank you for staying with us!", "Intent": "Transaction"}
Dataset Sources
Instances generated and annotated by ChatGPT 4.
Uses
Demo for email intent classification tasks.
email-summary-datasetturkish_spam_email
Dataset Info
Dataset obtained via https://www.kaggle.com/datasets/emrahaydemr/turkish-mail-dataset-normalspam
Urgency-tone-topic-on-enron_labeled_emails_with_subjects-llama2-7b_finetuningcustomer_care_emails
Dataset Card for customer_care_emails
This dataset contains synthetically generated emails that a customer care email unit will receive.
Dataset Details
Dataset Description
This dataset is a synthetically generated dataset using Gemini Pro. It is designed for the following hypothetical scenario.
Aetheros is a middleware solutions company for web apps. They have five main services: API development, API Monitoring, IAM, API development language called Mercury… See the full description on the dataset page: https://huggingface.co/datasets/rtweera/customer_care_emails.huggingface_filesystem_emails_terminal_6071_sales_orders_1787611667ner-email
Overview:
This dataset is augmented through llama-3.1:8b. The pourpose is to finetune llm for token classification i.e Email in our case.
Following tags are present in dataset:
full_name : 1
email : 2
gender : 3
city : 4
country : 5
legal-advice-email-risk-option-instruction-coherence-v0.1What this dataset does
You receive
case position
facts used
risk analysis
options
recommendation
client instruction
consistency flags
You decide
coherent
or
incoherent
Daily use
advice QC
risk gap detection
instruction capture check
contradiction flag
srilankan_email_datasetrestaurant-verified-email-access-in-columbus-ohio-us-173432
Restaurant Verified Email Access in Columbus, Ohio, US
Free sample dataset from BeamStation
Restaurant Verified Email Access in Columbus, Ohio, US
This dataset provides weekly‑verified email addresses for 665 established, independent restaurants (or micro‑chains) located in Columbus, Ohio. Large chain enterprises are excluded, ensuring the list focuses on independent operators. Each record includes a validated email address that has undergone our proprietary verification… See the full description on the dataset page: https://huggingface.co/datasets/beamstation/restaurant-verified-email-access-in-columbus-ohio-us-173432.spam-email-5k5
Spam Eail 5k5
DescriptionSpamMail-Binary is a curated email corpus designed for training and evaluating spam-detection systems. Each record contains:
Message – full email text, including subject and body
Category – binary label: Spam or Ham (non-spam)
The collection spans a diverse range of phishing attempts, promotional blasts, newsletters, and legitimate correspondence, offering clean, real-world language patterns for natural-language-processing and machine-learning tasks.
phishing-email-dataset
Phishing Email Dataset
This dataset on Hugging Face is a direct copy of the 'Phishing Email Detection' dataset from Kaggle, shared under the GNU Lesser General Public License 3.0. The dataset was originally created by the user 'Cyber Cop' on Kaggle. For complete details, including licensing and usage information, please visit the original Kaggle page.
email-intent-classificationCENSUS-NER-Name-Email-Address-PhoneDataset Summary
The CENSUS-NER-Name-Email-Address-Phone dataset is a processed and structured version of the FMCSA (Federal Motor Carrier Safety Administration) CENSUS1 2016Sep dataset. It is designed to assist in training language models for tasks such as Named Entity Recognition (NER), address parsing, and information extraction from unstructured text. The dataset contains records that include information such as name, email, phone number, and address, extracted from the original dataset and… See the full description on the dataset page: https://huggingface.co/datasets/Josephgflowers/CENSUS-NER-Name-Email-Address-Phone.german-english-email-ticket-classification
Customer Support Tickets (Short Version)
This dataset is a simplified version of the Customer Support Tickets dataset.
Dataset Details:
The dataset includes combinations of the following columns:
type
queue
priority
language
Modifications:
Shortened Version: This version only includes the first three rows for each combination of the above columns (i.e., 'type', 'queue', 'priority', 'language').
This reduction makes the dataset smaller and more manageable… See the full description on the dataset page: https://huggingface.co/datasets/ale-dp/german-english-email-ticket-classification.egal-client-instruction-email-call-note-action-coherence-risk-v0.1What this dataset does
You receive
instruction
channel
call note
action taken
confirmation sent
mismatch flags
You decide
coherent
or
incoherent
Daily use
instruction chain QC
“confirm in writing” enforcement
complaint risk reduction
enron_labeled_email-prompts-for-llama2_7b
