CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01corbt /enron-emailstext100K<n<1M13 likes25k downloads1y agoHugging Face02weaviate /enron-qa-emails-dasovich-jtext1K<n<10K0 likes2.3k downloads1y agoHugging Face03Matt1up /guertin-mcro-forensic-corpus-emails Guertin MCRO Forensic Corpus: Emails Contents: 352 email messages (.eml) in 11 categories — LinkedIn search-appearance notifications (99); messages before 2023-01-21 (82); messages after 2023-01-21 (76); correspondence with the first Rule 20 examiner (26); correspondence with the public defender (41); delivery of the March 5, 2025 hearing transcript (1); the Minnesota Attorney General's office (federal case) (1); 2026 correspondence (15); U.S. Senator Amy Klobuchar's office… See the full description on the dataset page: https://huggingface.co/datasets/Matt1up/guertin-mcro-forensic-corpus-emails.n<1K0 likes1.9k downloads11d agoHugging Face04snoop2head /enron_aeslc_emailstext100K<n<1M13 likes1.4k downloads4y agoHugging Face05from-our-page /hillary-clinton-emails-wikileakstext0 likes519 downloads1y agoHugging Face06corbt /enron_emails_sample_questionstabular10K<n<100K11 likes360 downloads10mo agoHugging Face07marketeam /Marketing-Emails Marketing Emails A curated corpus of synthetically generated yet realistic marketing email messages designed to support research in Domain Adaptation, Natural Language Processing (NLP), Data Science, Machine Learning, and Communication research. The dataset is appropriate for a wide spectrum of training paradigms—including pre-training, fine-tuning, and domain adaptation—as well as for rigorous evaluation of models targeting domain-specific language understanding and generation… See the full description on the dataset page: https://huggingface.co/datasets/marketeam/Marketing-Emails.texttext-generation10K<n<100K17 likes216 downloads10mo agoHugging Face08Hellisotherpeople /enron_emails_parsedtext100K<n<1M4 likes148 downloads3y agoHugging Face09hossein20s /enrun-emails-token-classificationtext10K<n<100K2 likes135 downloads4y agoHugging Face10UniqueData /email-spam-classification Email Spam Classification The dataset consists of a collection of emails categorized into two major classes: spam and not spam. It is designed to facilitate the development and evaluation of spam detection or email filtering systems. The spam emails in the dataset are typically unsolicited and unwanted messages that aim to promote products or services, spread malware, or deceive recipients for various malicious purposes. These emails often contain misleading subject lines… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/email-spam-classification.texttext-classificationn<1K10 likes131 downloads1y agoHugging Face11beyzasezer /marketing_emailstext0 likes130 downloads2y agoHugging Face12hossein20s /enrun-emails-text-classificationtext10K<n<100K5 likes129 downloads4y agoHugging Face13to-be /epstein-emails Epstein Email Messages Dataset Dataset Summary This dataset contains 4,272 individual email messages extracted directly from screenshot images using advanced Vision LLM technology. Unlike other datasets that work with only pre-OCR'd text, this dataset also used the original JPG screenshot images from the U.S. House Oversight Committee release using Qwen 2.5 VL 72B vision model to extract structured email data with high accuracy. The dataset powers a live Progressive Web… See the full description on the dataset page: https://huggingface.co/datasets/to-be/epstein-emails.tabulartext-classification1K<n<10K1 likes125 downloads10mo agoHugging Face14notesbymuneeb /epstein-emails Epstein Email Threads Dataset Dataset Summary This dataset contains 5,082 parsed email threads extracted from OCR'd documents released by the U.S. House Oversight Committee. The emails have been processed using large language models to extract structured information including senders, recipients, timestamps, subjects, and message bodies, with OCR errors corrected and footers removed. Dataset Description Overview This is a structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/notesbymuneeb/epstein-emails.texttext-classification1K<n<10K16 likes122 downloads10mo agoHugging Face15FaroukMoc2 /email_spam-qwen3-vl-32btext1K<n<10K0 likes96 downloads8mo agoHugging Face16jacquelinehe /enron-emailstext100K<n<1M0 likes94 downloads2y agoHugging Face17readerbench /ro-business-emailstext1K<n<10K2 likes73 downloads3y agoHugging Face18saberbx /Phishing_emails_testtextn<1K0 likes73 downloads1y agoHugging Face19drorrabin /phishing_emails-data 🛡️ Phishing Email Classification Dataset This dataset is curated for fine-tuning LLMs on the task of phishing email detection. It originates from this Kaggle dataset and has been transformed to better suit LLM-based classification tasks. 📦 Dataset Features Each row is a labeled email, with either: safe email (label = 0) phishing email (label = 1) The dataset includes metadata (sender, receiver, date, subject) and cleaned email body. Two main columns: Email Text:… See the full description on the dataset page: https://huggingface.co/datasets/drorrabin/phishing_emails-data.text10K<n<100K8 likes72 downloads1y agoHugging Face20ISTA-DASLab /Panza-emails The Panza Emails dataset This dataset contains collections of emails of three authentic users (david, isabel, and marcus), with personal information (names, places, etc.) replaced by other ones for donor privacy. Except for these changes, the language of the emails is genuine. The intention of this dataset is to allow researchers to study strategies for text personalization. The data was donated explicitly for this purpose. This dataset is ethically collected and fully licensed for… See the full description on the dataset page: https://huggingface.co/datasets/ISTA-DASLab/Panza-emails.textn<1K1 likes70 downloads2y agoHugging Face21NotShrirang /email-spam-filtertabulartext-classification1K<n<10K11 likes69 downloads1y agoHugging Face22Hannah2704 /epstein-emails Epstein Email Threads Dataset Dataset Summary This dataset contains 5,082 parsed email threads extracted from OCR'd documents released by the U.S. House Oversight Committee. The emails have been processed using large language models to extract structured information including senders, recipients, timestamps, subjects, and message bodies, with OCR errors corrected and footers removed. Dataset Description Overview This is a structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/Hannah2704/epstein-emails.texttext-classification1K<n<10K0 likes68 downloads8mo agoHugging Face23open-llm-leaderboard-old /details_postbot__pythia-160m-hq-emails Dataset Card for Evaluation run of postbot/pythia-160m-hq-emails Dataset Summary Dataset automatically created during the evaluation run of model postbot/pythia-160m-hq-emails on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_postbot__pythia-160m-hq-emails.0 likes66 downloads3y agoHugging Face24seuun /spam-ham-phish-emails-latesttext100K<n<1M0 likes64 downloads6mo agoHugging Face25TianfuXinqu /excel-huggingface-emails-1680-product-catalog Customer Satisfaction Survey 2024 Customer satisfaction survey responses collected from the mobile app between January and December 2024. Columns response_id customer_id satisfaction_score churn_risk survey_date textn<1K0 likes63 downloads1mo agoHugging Face26parameterlab /scaling_mia_the_pile_00_Enron_Emailstext10K<n<100K1 likes52 downloads2y agoHugging Face27timaeus /pile-enron_emails Dataset Creation Process These subsets were created by streaming over the rows from monology/pile-uncopyrighted and filtering by the meta column. Each subset is generally limited to the first 100,000 qualifying rows encountered. Citations If you use this dataset, please cite the original Pile papers: @article{gao2020pile, title={The Pile: An 800GB dataset of diverse text for language modeling}, author={Gao, Leo and Biderman, Stella and Black, Sid and Golding, Laurence and… See the full description on the dataset page: https://huggingface.co/datasets/timaeus/pile-enron_emails.text100K<n<1M2 likes50 downloads1y agoHugging Face28316usman /email-security EMAIL_SECURITY A preference dataset for EMAIL_SECURITY, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate. Format Standard preference / DPO schema — each row: column meaning prompt the request (originally prompt) chosen the human-preferred response rejected a worse response to the same prompt source the dataset/URL the row was harvested from Splits 80/10/10 train… See the full description on the dataset page: https://huggingface.co/datasets/316usman/email-security.texttext-generation1K<n<10K0 likes46 downloads9d agoHugging Face29Febriyansyah /phishing-emails-multilingual Phishing Emails Multilingual (ID/EN) — Synthetic Dataset sintetis & edukatif 600 email dwibahasa Indonesia 🇮🇩 & English 🇺🇸 untuk riset deteksi phishing — oleh Febriyansyah. ⚠️ Synthetic & edu-defense-only — dibuat untuk pembelajaran defensive security, bukan untuk kampanye nyata. Jangan gunakan untuk aktivitas ilegal. Ringkasan 600 baris — 300 phishing / 300 benign (seimbang), 321 EN / 279 ID Kolom: id (int), language (id/en), text (string, badan email)… See the full description on the dataset page: https://huggingface.co/datasets/Febriyansyah/phishing-emails-multilingual.tabulartext-classificationn<1K0 likes44 downloads18d agoHugging Face30weijianzhg /email-safety-triage-10k Email Safety Triage 10k This dataset contains 10,000 supervised examples for classifying email and email-adjacent content for operational triage, phishing/spam risk, and prompt-attack filtering. Each JSONL row has two string fields: input: an instruction plus email, security-review text, or prompt/email fragment. output: compact strict JSON with triage, priority, risk, should_process, confidence, and reason. The dataset is intended for fine-tuning and evaluating classifiers… See the full description on the dataset page: https://huggingface.co/datasets/weijianzhg/email-safety-triage-10k.texttext-classification10K<n<100K1 likes39 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.