datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FinePersonas-Synthetic-Email-Conversations
FinePersonas Synthetic Email Conversations
FinePersonas Synthetic Email Conversations is a dataset containing around 115k conversations via email between two personas from the argilla/FinePersonas-v0.1. Conversations were generated using NousResearch/Hermes-3-Llama-3.1-70B.
🗞️ News
[10/16/2024] New subsets: added two new subsets unfriendly_email_conversations and unprofessional_email_conversations.
How were the conversations generated?… See the full description on the dataset page: https://huggingface.co/datasets/argilla/FinePersonas-Synthetic-Email-Conversations.email-writing-sft-100k
Email Writing SFT (100K)
100,000 ShareGPT conversations demonstrating professional email writing across 22 business contexts. Each example shows how to draft clear, purposeful emails that achieve their communication goal — from cold outreach to salary negotiations to apology emails.
Motivation
Email is the primary communication channel for most professional work, yet LLMs often produce emails that are:
Too long: Including unnecessary preamble, excessive context… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/email-writing-sft-100k.Marketing-Emails
Marketing Emails
A curated corpus of synthetically generated yet realistic marketing email messages designed to support research in Domain Adaptation, Natural Language Processing (NLP), Data Science, Machine Learning, and Communication research.
The dataset is appropriate for a wide spectrum of training paradigms—including pre-training, fine-tuning, and domain adaptation—as well as for rigorous evaluation of models targeting domain-specific language understanding and generation… See the full description on the dataset page: https://huggingface.co/datasets/marketeam/Marketing-Emails.epstein-emails
Epstein Email Threads Dataset
Dataset Summary
This dataset contains 5,082 parsed email threads extracted from OCR'd documents released by the U.S. House Oversight Committee. The emails have been processed using large language models to extract structured information including senders, recipients, timestamps, subjects, and message bodies, with OCR errors corrected and footers removed.
Dataset Description
Overview
This is a structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/notesbymuneeb/epstein-emails.epstein-emails
Epstein Email Messages Dataset
Dataset Summary
This dataset contains 4,272 individual email messages extracted directly from screenshot images using advanced Vision LLM technology. Unlike other datasets that work with only pre-OCR'd text, this dataset also used the original JPG screenshot images from the U.S. House Oversight Committee release using Qwen 2.5 VL 72B vision model to extract structured email data with high accuracy.
The dataset powers a live Progressive Web… See the full description on the dataset page: https://huggingface.co/datasets/to-be/epstein-emails.Rail_Freight_Logistics_Company_Email_Archive_Sample
Ukrainian Rail-Freight Correspondence Corpus (Sample)
Real operational correspondence from a working freight forwarding business, and the
documents attached to it — consignment notes, service acts, invoices, wagon
manifests. Not scraped, not synthetic, and never published anywhere before.
This is a de-identified sample released for evaluation. It is drawn from a larger
private archive; see Full archive below.
Published by Akuma London · akumalondon.com
Why this… See the full description on the dataset page: https://huggingface.co/datasets/akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample.epstein-emails
Epstein Email Threads Dataset
Dataset Summary
This dataset contains 5,082 parsed email threads extracted from OCR'd documents released by the U.S. House Oversight Committee. The emails have been processed using large language models to extract structured information including senders, recipients, timestamps, subjects, and message bodies, with OCR errors corrected and footers removed.
Dataset Description
Overview
This is a structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/Hannah2704/epstein-emails.email-triage-v1
Email Triage v1
This dataset contains 1,740 unique email-triage examples produced through Tuned Tensor labeling and hardening workflows. It is the public v1 dataset, de-duplicated from a 2,038-row weighted fine-tuning dataset; 298 intentional weighting duplicates were removed for easier reuse.
The task is operational inbox triage, not security-risk classification. Each row asks a model to label one email-like message and return strict JSON with triage, priority, should_process… See the full description on the dataset page: https://huggingface.co/datasets/tunedtensor/email-triage-v1.business-email-dataset
Business Email Dataset - Alpaca Format
A comprehensive synthetic dataset of 5,000 professional business emails in Alpaca instruction-tuning format, designed for fine-tuning language models on formal business communication.
Dataset Description
This dataset contains high-quality, diverse business email examples covering a wide range of professional scenarios, industries, and communication styles. Each email is formatted following the Alpaca instruction-tuning standard… See the full description on the dataset page: https://huggingface.co/datasets/wardacoder/business-email-dataset.synthetic-abandoned-cart-email-examples
Synthetic Abandoned Cart Email Examples
An entirely synthetic, bilingual collection of abandoned-cart email drafts with transparent checklist annotations. It is intended for education, prototyping, and evaluation, and contains no real recipients, customer messages, orders, merchant data, or campaign results.
Dataset Description
The dataset mirrors the five visible checks in NeuroCheckout's public Abandoned Cart Email Checker:
message clarity;
primary call to… See the full description on the dataset page: https://huggingface.co/datasets/neurocheckout-ai/synthetic-abandoned-cart-email-examples.phishing-email-soc-agent
Phishing Email SOC Agent Dataset
A knowledge distillation dataset for training SOC (Security Operations Center) agents to detect and analyze phishing emails using tool-calling capabilities.
Dataset Description
This dataset contains 504 examples of email analysis with real tool calls and responses, designed for fine-tuning LLMs to become phishing detection agents. Each example includes:
Email parsing - Extract headers, URLs, IPs, attachments
Threat intelligence lookup -… See the full description on the dataset page: https://huggingface.co/datasets/Ellbendls/phishing-email-soc-agent.email-security
EMAIL_SECURITY
A preference dataset for EMAIL_SECURITY, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate.
Format
Standard preference / DPO schema — each row:
column
meaning
prompt
the request (originally prompt)
chosen
the human-preferred response
rejected
a worse response to the same prompt
source
the dataset/URL the row was harvested from
Splits
80/10/10 train… See the full description on the dataset page: https://huggingface.co/datasets/316usman/email-security.query-expansion
Query Expansion Dataset
This dataset is designed to train search query expansion models that can generate multiple semantic expansions for a given query.
Purpose
The goal of this dataset is to serve as input for training small language models (0.5B to 3B parameters) to act as query expander models in various search systems, including but not limited to Retrieval-Augmented Generation (RAG) systems.
Query expansion is a technique used to enhance search results by generating… See the full description on the dataset page: https://huggingface.co/datasets/s-emanuilov/query-expansion.pleias-post-ocr-correction-chonkie-aligned-en
PleIAs Post-OCR Correction — Chonkie-Aligned Semantic Chunks
This dataset is a semantically chunked and span-aligned derivative of PleIAs/Post-OCR-Correction.
Each record contains:
an OCR hypothesis chunk from the original text field;
a corresponding post-OCR correction output chunk from the corrected_text field;
metadata inherited from the PleIAs dataset;
character spans linking each chunk back to the original source document;
alignment diagnostics produced during filtering.… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/pleias-post-ocr-correction-chonkie-aligned-en.email-safety-triage-10k
Email Safety Triage 10k
This dataset contains 10,000 supervised examples for classifying email and email-adjacent content for operational triage, phishing/spam risk, and prompt-attack filtering.
Each JSONL row has two string fields:
input: an instruction plus email, security-review text, or prompt/email fragment.
output: compact strict JSON with triage, priority, risk, should_process, confidence, and reason.
The dataset is intended for fine-tuning and evaluating classifiers… See the full description on the dataset page: https://huggingface.co/datasets/weijianzhg/email-safety-triage-10k.em-activation-transport
EM activation transport — paired activations, generations, and judgments
Research artifacts from two experiments on inference-time activation transport for
emergent misalignment (EM), using the released model organism
ModelOrganismsForEM/Llama-3.1-8B-Instruct_bad-medical-advice
(rank-32 LoRA) against clean
meta-llama/Llama-3.1-8B-Instruct.
Why this repo exists: activations and generations are the expensive, fixed part of
these experiments; judgments are cheap and revisable. In… See the full description on the dataset page: https://huggingface.co/datasets/assisted-diffusion/em-activation-transport.Jeffrey-Epstein-Emails-From-Epstein-Files
Jeffrey Epstein Emails from Epstein Files
This dataset contains 7,380 emails with 14,835 messages scraped from jmail.world.
Dataset Description
This dataset provides email correspondence from the Jeffrey Epstein Files, scraped directly from the jmail.world website.
Data Source
All emails were scraped from the jmail.world website.
Fields
Field
Type
Description
doc_id
string
Internal document identifier
subject
string
Email subject line… See the full description on the dataset page: https://huggingface.co/datasets/KillerShoaib/Jeffrey-Epstein-Emails-From-Epstein-Files.email-datasets-20k
Dataset Summary
There are 20,000 samples of emails.
This dataset was created using Gemma 3-4B-it (via mlx-community/gemma-3-4b-it-4bit-DWQ).
License Note
This dataset is licensed under Apache 2.0. Please also refer to the Gemma Terms of Use and Prohibited Use Policy regarding the use of Gemma-generated content.
chrononoise-claims-fr
ChronoNoise-Claims-FR
ChronoNoise-Claims-FR is a silver dataset for studying whether LLM-generated historical claims are supported by noisy OCR text, post-corrected text, and document metadata.
The dataset is designed around a common historical NLP failure chain:
historical OCR noise → fluent LLM interpretation → plausible but unsupported historical claim
It is derived from a ChronoCorrect-Europeana-style dataset built from historical French newspaper OCR. Each record contains… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/chrononoise-claims-fr.pleias-post-ocr-correction-chonkie-aligned-fr
PleIAs Post-OCR Correction — Chonkie-Aligned Semantic Chunks
This dataset is a semantically chunked and span-aligned derivative of PleIAs/Post-OCR-Correction.
Each record contains:
an OCR hypothesis chunk from the original text field;
a corresponding post-OCR correction output chunk from the corrected_text field;
metadata inherited from the PleIAs dataset;
character spans linking each chunk back to the original source document;
alignment diagnostics produced during filtering.… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/pleias-post-ocr-correction-chonkie-aligned-fr.epstein-emails
Epstein Email Threads Dataset
Dataset Summary
This dataset contains 5,082 parsed email threads extracted from OCR'd documents released by the U.S. House Oversight Committee. The emails have been processed using large language models to extract structured information including senders, recipients, timestamps, subjects, and message bodies, with OCR errors corrected and footers removed.
Dataset Description
Overview
This is a structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/pupepps/epstein-emails.enron_labeled_email-prompts-for-llama2_7bemail-datasets-v2-100k
Dataset Summary
There are 99336 samples of emails.
This dataset was created using Gemma 3-4B-it (via mlx-community/gemma-3-4b-it-4bit-DWQ).
format:
{"id": , "instruction": "Prompt Is Here",
"text":
"<user>Prompt Is Here
<think>
- Goal: Goal Is Here
- Reason: Reason Is Here
- Tone: Tone Is Here
</think>
<generate>
Mail Is Here
</generate></s>"}
Link
Github: https://github.com/kamisori-daijin/email-datasets
License Note
This dataset is licensed… See the full description on the dataset page: https://huggingface.co/datasets/Kamisori-daijin/email-datasets-v2-100k.Emakhuwa-Portuguese-OCR-post-correctionBibTeX:
The dataset paper was published in EMNLP 2024.
Please cite as:
@inproceedings{ali-etal-2024-building,
title = "Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks",
author = "Ali, Felermino D. M. A. and
Lopes Cardoso, Henrique and
Sousa-Silva, Rui",
editor = "Al-Onaizan, Yaser and
Bansal, Mohit and
Chen, Yun-Nung",
booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-Portuguese-OCR-post-correction.MatLab-Questions-Answers
Dataset Card for Dataset Name
Matlab-Questions-Answers
Dataset Details
Contains 150+ matlab/octave related questions and answers.
Dataset Description
Contains 150+ matlab/octave related questions and answers ranging from Grade School to Graduate level mathematics.
Language(s) (NLP): English
License: Apache 2.0
Uses
Small Language Models on matlab/octave specific code generation
Evaluation of Language Models on Matlab related questions and answers… See the full description on the dataset page: https://huggingface.co/datasets/EmAllen-TW/MatLab-Questions-Answers.Emakhuwa-MonolingualBibTeX:
The dataset paper was published in EMNLP 2024.
Please cite as:
@inproceedings{ali-etal-2024-building,
title = "Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks",
author = "Ali, Felermino D. M. A. and
Lopes Cardoso, Henrique and
Sousa-Silva, Rui",
editor = "Al-Onaizan, Yaser and
Bansal, Mohit and
Chen, Yun-Nung",
booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-Monolingual.smolified-sample-email-generator
🤏 smolified-sample-email-generator
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-sample-email-generator.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 6e8b8fbf)
Records: 244
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
EmailBot2pashto-alpaca-business-emails
📧 Pashto Alpaca Business Emails Dataset
د پښتو سوداګریز بریښنالیکونو ډیټاسیټ
🌟 د سوداګریزو بریښنالیکونو لپاره تر ټولو لوی پښتو ډیټاسیټThe largest Pashto dataset for business email generation
This is a meticulously curated, high-quality dataset of business emails translated into Pashto, designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) for professional communication in Pashto.
🎯 Why This Dataset Matters
Challenge
Our… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-alpaca-business-emails.smolified-email-tone-converter
🤏 smolified-email-tone-converter
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model ankitaroy07/smolified-email-tone-converter.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 93e3ec46)
Records: 1480
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by ankitaroy07.
Generated via Smolify.ai.
