CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01argilla /FinePersonas-Synthetic-Email-Conversations FinePersonas Synthetic Email Conversations FinePersonas Synthetic Email Conversations is a dataset containing around 115k conversations via email between two personas from the argilla/FinePersonas-v0.1. Conversations were generated using NousResearch/Hermes-3-Llama-3.1-70B. 🗞️ News [10/16/2024] New subsets: added two new subsets unfriendly_email_conversations and unprofessional_email_conversations. How were the conversations generated?… See the full description on the dataset page: https://huggingface.co/datasets/argilla/FinePersonas-Synthetic-Email-Conversations.texttext-generation100K<n<1M8 likes399 downloads2y agoHugging Face02stindardlogic /email-writing-sft-100k Email Writing SFT (100K) 100,000 ShareGPT conversations demonstrating professional email writing across 22 business contexts. Each example shows how to draft clear, purposeful emails that achieve their communication goal — from cold outreach to salary negotiations to apology emails. Motivation Email is the primary communication channel for most professional work, yet LLMs often produce emails that are: Too long: Including unnecessary preamble, excessive context… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/email-writing-sft-100k.texttext-generation100K<n<1M3 likes243 downloads2mo agoHugging Face03marketeam /Marketing-Emails Marketing Emails A curated corpus of synthetically generated yet realistic marketing email messages designed to support research in Domain Adaptation, Natural Language Processing (NLP), Data Science, Machine Learning, and Communication research. The dataset is appropriate for a wide spectrum of training paradigms—including pre-training, fine-tuning, and domain adaptation—as well as for rigorous evaluation of models targeting domain-specific language understanding and generation… See the full description on the dataset page: https://huggingface.co/datasets/marketeam/Marketing-Emails.texttext-generation10K<n<100K17 likes202 downloads10mo agoHugging Face04notesbymuneeb /epstein-emails Epstein Email Threads Dataset Dataset Summary This dataset contains 5,082 parsed email threads extracted from OCR'd documents released by the U.S. House Oversight Committee. The emails have been processed using large language models to extract structured information including senders, recipients, timestamps, subjects, and message bodies, with OCR errors corrected and footers removed. Dataset Description Overview This is a structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/notesbymuneeb/epstein-emails.texttext-classification1K<n<10K16 likes123 downloads10mo agoHugging Face05to-be /epstein-emails Epstein Email Messages Dataset Dataset Summary This dataset contains 4,272 individual email messages extracted directly from screenshot images using advanced Vision LLM technology. Unlike other datasets that work with only pre-OCR'd text, this dataset also used the original JPG screenshot images from the U.S. House Oversight Committee release using Qwen 2.5 VL 72B vision model to extract structured email data with high accuracy. The dataset powers a live Progressive Web… See the full description on the dataset page: https://huggingface.co/datasets/to-be/epstein-emails.tabulartext-classification1K<n<10K1 likes121 downloads10mo agoHugging Face06akumalondon /Rail_Freight_Logistics_Company_Email_Archive_Sample Ukrainian Rail-Freight Correspondence Corpus (Sample) Real operational correspondence from a working freight forwarding business, and the documents attached to it — consignment notes, service acts, invoices, wagon manifests. Not scraped, not synthetic, and never published anywhere before. This is a de-identified sample released for evaluation. It is drawn from a larger private archive; see Full archive below. Published by Akuma London · akumalondon.com Why this… See the full description on the dataset page: https://huggingface.co/datasets/akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample.tabulartext-generation1K<n<10K0 likes90 downloads12d agoHugging Face07Hannah2704 /epstein-emails Epstein Email Threads Dataset Dataset Summary This dataset contains 5,082 parsed email threads extracted from OCR'd documents released by the U.S. House Oversight Committee. The emails have been processed using large language models to extract structured information including senders, recipients, timestamps, subjects, and message bodies, with OCR errors corrected and footers removed. Dataset Description Overview This is a structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/Hannah2704/epstein-emails.texttext-classification1K<n<10K0 likes67 downloads8mo agoHugging Face08tunedtensor /email-triage-v1 Email Triage v1 This dataset contains 1,740 unique email-triage examples produced through Tuned Tensor labeling and hardening workflows. It is the public v1 dataset, de-duplicated from a 2,038-row weighted fine-tuning dataset; 298 intentional weighting duplicates were removed for easier reuse. The task is operational inbox triage, not security-risk classification. Each row asks a model to label one email-like message and return strict JSON with triage, priority, should_process… See the full description on the dataset page: https://huggingface.co/datasets/tunedtensor/email-triage-v1.tabulartext-classification1K<n<10K0 likes64 downloads3mo agoHugging Face09wardacoder /business-email-dataset Business Email Dataset - Alpaca Format A comprehensive synthetic dataset of 5,000 professional business emails in Alpaca instruction-tuning format, designed for fine-tuning language models on formal business communication. Dataset Description This dataset contains high-quality, diverse business email examples covering a wide range of professional scenarios, industries, and communication styles. Each email is formatted following the Alpaca instruction-tuning standard… See the full description on the dataset page: https://huggingface.co/datasets/wardacoder/business-email-dataset.texttext-generation10K<n<100K0 likes61 downloads1y agoHugging Face10neurocheckout-ai /synthetic-abandoned-cart-email-examples Synthetic Abandoned Cart Email Examples An entirely synthetic, bilingual collection of abandoned-cart email drafts with transparent checklist annotations. It is intended for education, prototyping, and evaluation, and contains no real recipients, customer messages, orders, merchant data, or campaign results. Dataset Description The dataset mirrors the five visible checks in NeuroCheckout's public Abandoned Cart Email Checker: message clarity; primary call to… See the full description on the dataset page: https://huggingface.co/datasets/neurocheckout-ai/synthetic-abandoned-cart-email-examples.tabulartext-classificationn<1K0 likes60 downloads24d agoHugging Face11Ellbendls /phishing-email-soc-agent Phishing Email SOC Agent Dataset A knowledge distillation dataset for training SOC (Security Operations Center) agents to detect and analyze phishing emails using tool-calling capabilities. Dataset Description This dataset contains 504 examples of email analysis with real tool calls and responses, designed for fine-tuning LLMs to become phishing detection agents. Each example includes: Email parsing - Extract headers, URLs, IPs, attachments Threat intelligence lookup -… See the full description on the dataset page: https://huggingface.co/datasets/Ellbendls/phishing-email-soc-agent.texttext-generationn<1K0 likes46 downloads7mo agoHugging Face12316usman /email-security EMAIL_SECURITY A preference dataset for EMAIL_SECURITY, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate. Format Standard preference / DPO schema — each row: column meaning prompt the request (originally prompt) chosen the human-preferred response rejected a worse response to the same prompt source the dataset/URL the row was harvested from Splits 80/10/10 train… See the full description on the dataset page: https://huggingface.co/datasets/316usman/email-security.texttext-generation1K<n<10K0 likes46 downloads8d agoHugging Face13s-emanuilov /query-expansion Query Expansion Dataset This dataset is designed to train search query expansion models that can generate multiple semantic expansions for a given query. Purpose The goal of this dataset is to serve as input for training small language models (0.5B to 3B parameters) to act as query expander models in various search systems, including but not limited to Retrieval-Augmented Generation (RAG) systems. Query expansion is a technique used to enhance search results by generating… See the full description on the dataset page: https://huggingface.co/datasets/s-emanuilov/query-expansion.texttext-generation1K<n<10K4 likes45 downloads2y agoHugging Face14emanuelaboros /pleias-post-ocr-correction-chonkie-aligned-en PleIAs Post-OCR Correction — Chonkie-Aligned Semantic Chunks This dataset is a semantically chunked and span-aligned derivative of PleIAs/Post-OCR-Correction. Each record contains: an OCR hypothesis chunk from the original text field; a corresponding post-OCR correction output chunk from the corrected_text field; metadata inherited from the PleIAs dataset; character spans linking each chunk back to the original source document; alignment diagnostics produced during filtering.… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/pleias-post-ocr-correction-chonkie-aligned-en.texttext-generation100K<n<1M0 likes40 downloads3mo agoHugging Face15weijianzhg /email-safety-triage-10k Email Safety Triage 10k This dataset contains 10,000 supervised examples for classifying email and email-adjacent content for operational triage, phishing/spam risk, and prompt-attack filtering. Each JSONL row has two string fields: input: an instruction plus email, security-review text, or prompt/email fragment. output: compact strict JSON with triage, priority, risk, should_process, confidence, and reason. The dataset is intended for fine-tuning and evaluating classifiers… See the full description on the dataset page: https://huggingface.co/datasets/weijianzhg/email-safety-triage-10k.texttext-classification10K<n<100K1 likes38 downloads4mo agoHugging Face16assisted-diffusion /em-activation-transport EM activation transport — paired activations, generations, and judgments Research artifacts from two experiments on inference-time activation transport for emergent misalignment (EM), using the released model organism ModelOrganismsForEM/Llama-3.1-8B-Instruct_bad-medical-advice (rank-32 LoRA) against clean meta-llama/Llama-3.1-8B-Instruct. Why this repo exists: activations and generations are the expensive, fixed part of these experiments; judgments are cheap and revisable. In… See the full description on the dataset page: https://huggingface.co/datasets/assisted-diffusion/em-activation-transport.texttext-generationn<1K0 likes37 downloads2mo agoHugging Face17KillerShoaib /Jeffrey-Epstein-Emails-From-Epstein-Files Jeffrey Epstein Emails from Epstein Files This dataset contains 7,380 emails with 14,835 messages scraped from jmail.world. Dataset Description This dataset provides email correspondence from the Jeffrey Epstein Files, scraped directly from the jmail.world website. Data Source All emails were scraped from the jmail.world website. Fields Field Type Description doc_id string Internal document identifier subject string Email subject line… See the full description on the dataset page: https://huggingface.co/datasets/KillerShoaib/Jeffrey-Epstein-Emails-From-Epstein-Files.texttext-classification10K<n<100K4 likes29 downloads7mo agoHugging Face18Kamisori-daijin /email-datasets-20k Dataset Summary There are 20,000 samples of emails. This dataset was created using Gemma 3-4B-it (via mlx-community/gemma-3-4b-it-4bit-DWQ). License Note This dataset is licensed under Apache 2.0. Please also refer to the Gemma Terms of Use and Prohibited Use Policy regarding the use of Gemma-generated content. texttext-generation10K<n<100K3 likes28 downloads6mo agoHugging Face19emanuelaboros /chrononoise-claims-fr ChronoNoise-Claims-FR ChronoNoise-Claims-FR is a silver dataset for studying whether LLM-generated historical claims are supported by noisy OCR text, post-corrected text, and document metadata. The dataset is designed around a common historical NLP failure chain: historical OCR noise → fluent LLM interpretation → plausible but unsupported historical claim It is derived from a ChronoCorrect-Europeana-style dataset built from historical French newspaper OCR. Each record contains… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/chrononoise-claims-fr.texttext-classificationn<1K0 likes27 downloads3mo agoHugging Face20emanuelaboros /pleias-post-ocr-correction-chonkie-aligned-fr PleIAs Post-OCR Correction — Chonkie-Aligned Semantic Chunks This dataset is a semantically chunked and span-aligned derivative of PleIAs/Post-OCR-Correction. Each record contains: an OCR hypothesis chunk from the original text field; a corresponding post-OCR correction output chunk from the corrected_text field; metadata inherited from the PleIAs dataset; character spans linking each chunk back to the original source document; alignment diagnostics produced during filtering.… See the full description on the dataset page: https://huggingface.co/datasets/emanuelaboros/pleias-post-ocr-correction-chonkie-aligned-fr.texttext-generation10K<n<100K0 likes24 downloads3mo agoHugging Face21pupepps /epstein-emails Epstein Email Threads Dataset Dataset Summary This dataset contains 5,082 parsed email threads extracted from OCR'd documents released by the U.S. House Oversight Committee. The emails have been processed using large language models to extract structured information including senders, recipients, timestamps, subjects, and message bodies, with OCR errors corrected and footers removed. Dataset Description Overview This is a structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/pupepps/epstein-emails.texttext-classification1K<n<10K1 likes23 downloads8mo agoHugging Face22cheekymachine /enron_labeled_email-prompts-for-llama2_7btexttext-classification1K<n<10K0 likes18 downloads3y agoHugging Face23Kamisori-daijin /email-datasets-v2-100k Dataset Summary There are 99336 samples of emails. This dataset was created using Gemma 3-4B-it (via mlx-community/gemma-3-4b-it-4bit-DWQ). format: {"id": , "instruction": "Prompt Is Here", "text": "<user>Prompt Is Here <think> - Goal: Goal Is Here - Reason: Reason Is Here - Tone: Tone Is Here </think> <generate> Mail Is Here </generate></s>"} Link Github: https://github.com/kamisori-daijin/email-datasets License Note This dataset is licensed… See the full description on the dataset page: https://huggingface.co/datasets/Kamisori-daijin/email-datasets-v2-100k.texttext-generation10K<n<100K0 likes17 downloads5mo agoHugging Face24LIACC /Emakhuwa-Portuguese-OCR-post-correctionBibTeX: The dataset paper was published in EMNLP 2024. Please cite as: @inproceedings{ali-etal-2024-building, title = "Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks", author = "Ali, Felermino D. M. A. and Lopes Cardoso, Henrique and Sousa-Silva, Rui", editor = "Al-Onaizan, Yaser and Bansal, Mohit and Chen, Yun-Nung", booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-Portuguese-OCR-post-correction.imagetranslationn<1K0 likes16 downloads2y agoHugging Face25EmAllen-TW /MatLab-Questions-Answers Dataset Card for Dataset Name Matlab-Questions-Answers Dataset Details Contains 150+ matlab/octave related questions and answers. Dataset Description Contains 150+ matlab/octave related questions and answers ranging from Grade School to Graduate level mathematics. Language(s) (NLP): English License: Apache 2.0 Uses Small Language Models on matlab/octave specific code generation Evaluation of Language Models on Matlab related questions and answers… See the full description on the dataset page: https://huggingface.co/datasets/EmAllen-TW/MatLab-Questions-Answers.texttext-generationn<1K0 likes13 downloads1y agoHugging Face26LIACC /Emakhuwa-MonolingualBibTeX: The dataset paper was published in EMNLP 2024. Please cite as: @inproceedings{ali-etal-2024-building, title = "Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks", author = "Ali, Felermino D. M. A. and Lopes Cardoso, Henrique and Sousa-Silva, Rui", editor = "Al-Onaizan, Yaser and Bansal, Mohit and Chen, Yun-Nung", booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-Monolingual.tabulartranslation10K<n<100K0 likes11 downloads2y agoHugging Face27smolify /smolified-sample-email-generator 🤏 smolified-sample-email-generator Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model smolify/smolified-sample-email-generator. 📦 Asset Details Origin: Smolify Foundry (Job ID: 6e8b8fbf) Records: 244 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by smolify. Generated via Smolify.ai. texttext-generationn<1K0 likes9 downloads6mo agoHugging Face28carseng /EmailBot2gatedtexttext-generationn<1K0 likes8 downloads9mo agoHugging Face29nassimjp /pashto-alpaca-business-emails 📧 Pashto Alpaca Business Emails Dataset د پښتو سوداګریز بریښنالیکونو ډیټاسیټ 🌟 د سوداګریزو بریښنالیکونو لپاره تر ټولو لوی پښتو ډیټاسیټThe largest Pashto dataset for business email generation This is a meticulously curated, high-quality dataset of business emails translated into Pashto, designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) for professional communication in Pashto. 🎯 Why This Dataset Matters Challenge Our… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-alpaca-business-emails.texttext-generation1K<n<10K0 likes8 downloads4mo agoHugging Face30ankitaroy07 /smolified-email-tone-converter 🤏 smolified-email-tone-converter Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model ankitaroy07/smolified-email-tone-converter. 📦 Asset Details Origin: Smolify Foundry (Job ID: 93e3ec46) Records: 1480 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by ankitaroy07. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes7 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.