datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Marketing-Emails
Marketing Emails
A curated corpus of synthetically generated yet realistic marketing email messages designed to support research in Domain Adaptation, Natural Language Processing (NLP), Data Science, Machine Learning, and Communication research.
The dataset is appropriate for a wide spectrum of training paradigms—including pre-training, fine-tuning, and domain adaptation—as well as for rigorous evaluation of models targeting domain-specific language understanding and generation… See the full description on the dataset page: https://huggingface.co/datasets/marketeam/Marketing-Emails.epstein-emails
Epstein Email Messages Dataset
Dataset Summary
This dataset contains 4,272 individual email messages extracted directly from screenshot images using advanced Vision LLM technology. Unlike other datasets that work with only pre-OCR'd text, this dataset also used the original JPG screenshot images from the U.S. House Oversight Committee release using Qwen 2.5 VL 72B vision model to extract structured email data with high accuracy.
The dataset powers a live Progressive Web… See the full description on the dataset page: https://huggingface.co/datasets/to-be/epstein-emails.epstein-emails
Epstein Email Threads Dataset
Dataset Summary
This dataset contains 5,082 parsed email threads extracted from OCR'd documents released by the U.S. House Oversight Committee. The emails have been processed using large language models to extract structured information including senders, recipients, timestamps, subjects, and message bodies, with OCR errors corrected and footers removed.
Dataset Description
Overview
This is a structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/notesbymuneeb/epstein-emails.epstein-emails
Epstein Email Threads Dataset
Dataset Summary
This dataset contains 5,082 parsed email threads extracted from OCR'd documents released by the U.S. House Oversight Committee. The emails have been processed using large language models to extract structured information including senders, recipients, timestamps, subjects, and message bodies, with OCR errors corrected and footers removed.
Dataset Description
Overview
This is a structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/Hannah2704/epstein-emails.email-security
EMAIL_SECURITY
A preference dataset for EMAIL_SECURITY, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate.
Format
Standard preference / DPO schema — each row:
column
meaning
prompt
the request (originally prompt)
chosen
the human-preferred response
rejected
a worse response to the same prompt
source
the dataset/URL the row was harvested from
Splits
80/10/10 train… See the full description on the dataset page: https://huggingface.co/datasets/316usman/email-security.email-safety-triage-10k
Email Safety Triage 10k
This dataset contains 10,000 supervised examples for classifying email and email-adjacent content for operational triage, phishing/spam risk, and prompt-attack filtering.
Each JSONL row has two string fields:
input: an instruction plus email, security-review text, or prompt/email fragment.
output: compact strict JSON with triage, priority, risk, should_process, confidence, and reason.
The dataset is intended for fine-tuning and evaluating classifiers… See the full description on the dataset page: https://huggingface.co/datasets/weijianzhg/email-safety-triage-10k.Jeffrey-Epstein-Emails-From-Epstein-Files
Jeffrey Epstein Emails from Epstein Files
This dataset contains 7,380 emails with 14,835 messages scraped from jmail.world.
Dataset Description
This dataset provides email correspondence from the Jeffrey Epstein Files, scraped directly from the jmail.world website.
Data Source
All emails were scraped from the jmail.world website.
Fields
Field
Type
Description
doc_id
string
Internal document identifier
subject
string
Email subject line… See the full description on the dataset page: https://huggingface.co/datasets/KillerShoaib/Jeffrey-Epstein-Emails-From-Epstein-Files.epstein-emails
Epstein Email Threads Dataset
Dataset Summary
This dataset contains 5,082 parsed email threads extracted from OCR'd documents released by the U.S. House Oversight Committee. The emails have been processed using large language models to extract structured information including senders, recipients, timestamps, subjects, and message bodies, with OCR errors corrected and footers removed.
Dataset Description
Overview
This is a structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/pupepps/epstein-emails.pashto-alpaca-business-emails
📧 Pashto Alpaca Business Emails Dataset
د پښتو سوداګریز بریښنالیکونو ډیټاسیټ
🌟 د سوداګریزو بریښنالیکونو لپاره تر ټولو لوی پښتو ډیټاسیټThe largest Pashto dataset for business email generation
This is a meticulously curated, high-quality dataset of business emails translated into Pashto, designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) for professional communication in Pashto.
🎯 Why This Dataset Matters
Challenge
Our… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-alpaca-business-emails.
