datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
enron-emailsenron-qa-emails-dasovich-jfeatures-dinov3-vith16plus-224-imagenet-22k-wdsenron-emailThis dataset includes emails from Enron Email Dataset with prompts processed from Are Large Pre-Trained Language Models Leaking Your Personal Information?.
To use the dataset, you can run the following in LLM-PBE.
from data.enron import EnronDataset
ds = EnronDataset(data_path="data/enron", pseudonymize=False)
enron_aeslc_emailsphishing-email-dataset
Phishing Email Dataset
This dataset on Hugging Face is a direct copy of the 'Phishing Email Detection' dataset from Kaggle, shared under the GNU Lesser General Public License 3.0. The dataset was originally created by the user 'Cyber Cop' on Kaggle. For complete details, including licensing and usage information, please visit the original Kaggle page.
seven-phishing-email-datasets
Dataset Card for Seven Phishing/Spam Email Datasets
Dataset Summary
This dataset is a unified, row-level email corpus built from seven commonly used public email datasets. It is intended for research on phishing/spam detection and related email-text classification tasks.
Each row contains the email body (text), optional header-like fields (e.g., sender, receiver, date), the source dataset name (dataset_name), and a binary label (label).
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/puyang2025/seven-phishing-email-datasets.enron_emailshadow-eo
Dataset Card for S-EO: A Large-Scale Dataset for Geometry-Aware Shadow Detection in Remote Sensing Applications
Project page
We introduce the S-EO dataset: a large-scale, high-resolution dataset designed to advance geometry-aware shadow detection. Collected from diverse public-domain sources, including challenge datasets and government providers such as USGS, our dataset comprises 702 georeferenced tiles across the USA, each covering 500 × 500 meters. Each tile includes multi-date… See the full description on the dataset page: https://huggingface.co/datasets/emasquil/shadow-eo.hillary-clinton-emails-wikileaksvertebrate_genomesenron_emails_sample_questionsFinePersonas-Synthetic-Email-Conversations
FinePersonas Synthetic Email Conversations
FinePersonas Synthetic Email Conversations is a dataset containing around 115k conversations via email between two personas from the argilla/FinePersonas-v0.1. Conversations were generated using NousResearch/Hermes-3-Llama-3.1-70B.
🗞️ News
[10/16/2024] New subsets: added two new subsets unfriendly_email_conversations and unprofessional_email_conversations.
How were the conversations generated?… See the full description on the dataset page: https://huggingface.co/datasets/argilla/FinePersonas-Synthetic-Email-Conversations.small_vertebrate_genomes_8192train5
Dataset Card for "train5"
More Information needed
email-writing-sft-100k
Email Writing SFT (100K)
100,000 ShareGPT conversations demonstrating professional email writing across 22 business contexts. Each example shows how to draft clear, purposeful emails that achieve their communication goal — from cold outreach to salary negotiations to apology emails.
Motivation
Email is the primary communication channel for most professional work, yet LLMs often produce emails that are:
Too long: Including unnecessary preamble, excessive context… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/email-writing-sft-100k.cls_dinov3-vith16plus_in22ktrain4
Dataset Card for "train4"
More Information needed
train3
Dataset Card for "train3"
More Information needed
train1
Dataset Card for "train1"
More Information needed
Marketing-Emails
Marketing Emails
A curated corpus of synthetically generated yet realistic marketing email messages designed to support research in Domain Adaptation, Natural Language Processing (NLP), Data Science, Machine Learning, and Communication research.
The dataset is appropriate for a wide spectrum of training paradigms—including pre-training, fine-tuning, and domain adaptation—as well as for rigorous evaluation of models targeting domain-specific language understanding and generation… See the full description on the dataset page: https://huggingface.co/datasets/marketeam/Marketing-Emails.train6
Dataset Card for "train6"
More Information needed
train2
Dataset Card for "train2"
More Information needed
train8
Dataset Card for "train8"
More Information needed
train7
Dataset Card for "train7"
More Information needed
phish-email-datasets
Dataset Card: Phish Email Datasets
Dataset Summary
This repository contains three Parquet files of email records for phishing-related modeling and analysis.
Total rows across files: 7,962
Common fields: sender/receiver metadata, date string, subject, body, URL indicator or URL content, label
Storage format: Apache Parquet
Repository Contents
File
Rows
Columns
Label distribution
Nazario.parquet
1,565
7
{1: 1565}
Nazario_5.parquet
3,065
7
{1: 1565… See the full description on the dataset page: https://huggingface.co/datasets/puyang2025/phish-email-datasets.enron_emails_parsedAngiosperm_65_genomes_8192bp_rcemail-thread-summary
Dataset Card for "email-thread-summary"
More Information needed
mammals_32k
