CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zefang-liu /phishing-email-dataset Phishing Email Dataset This dataset on Hugging Face is a direct copy of the 'Phishing Email Detection' dataset from Kaggle, shared under the GNU Lesser General Public License 3.0. The dataset was originally created by the user 'Cyber Cop' on Kaggle. For complete details, including licensing and usage information, please visit the original Kaggle page. texttext-classification10K<n<100K38 likes1.2k downloads3y agoHugging Face02marketeam /Marketing-Emails Marketing Emails A curated corpus of synthetically generated yet realistic marketing email messages designed to support research in Domain Adaptation, Natural Language Processing (NLP), Data Science, Machine Learning, and Communication research. The dataset is appropriate for a wide spectrum of training paradigms—including pre-training, fine-tuning, and domain adaptation—as well as for rigorous evaluation of models targeting domain-specific language understanding and generation… See the full description on the dataset page: https://huggingface.co/datasets/marketeam/Marketing-Emails.texttext-generation10K<n<100K17 likes200 downloads10mo agoHugging Face03UniqueData /email-spam-classification Email Spam Classification The dataset consists of a collection of emails categorized into two major classes: spam and not spam. It is designed to facilitate the development and evaluation of spam detection or email filtering systems. The spam emails in the dataset are typically unsolicited and unwanted messages that aim to promote products or services, spread malware, or deceive recipients for various malicious purposes. These emails often contain misleading subject lines… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/email-spam-classification.texttext-classificationn<1K10 likes129 downloads1y agoHugging Face04locuoco /the-biggest-spam-ham-phish-email-dataset-300000 The Biggest Spam Ham Phish Email Dataset (250000+) This dataset is a large-scale, unified, and deduplicated collection of text messages and emails created for spam, ham, and phishing detection. It has been constructed by combining multiple publicly available and open-source datasets into a single standardized format, making it suitable for machine learning, deep learning, and NLP-based projects. The dataset contains approximately unique 250,000+ samples, covering a diverse range of… See the full description on the dataset page: https://huggingface.co/datasets/locuoco/the-biggest-spam-ham-phish-email-dataset-300000.texttext-classification100K<n<1M0 likes116 downloads7mo agoHugging Face05Dizzzy0x00 /LLMGen-Phishing-Email-Dataset LLM-Generated Phishing Email Dataset Dataset Description This dataset comprises a collection of phishing and legitimate emails generated using Large Language Models (LLMs), specifically DeepSeek for Chinese emails and OpenAI models for English emails. The primary purpose of this dataset is to facilitate research and development in phishing email detection and classification. The dataset is structured with two key columns: content: The full text content of the email.… See the full description on the dataset page: https://huggingface.co/datasets/Dizzzy0x00/LLMGen-Phishing-Email-Dataset.texttext-classification1K<n<10K2 likes107 downloads10mo agoHugging Face06jhonrayo99 /phishing-email-balanced-6000 Balanced Phishing Email Detection Subset This dataset is a derived, randomly sampled subset of Cyber Cop's Phishing Email Detection dataset on Kaggle. The original dataset is distributed under the GNU Lesser General Public License 3.0. Dataset structure The file phishing_email_subset.csv contains 6,000 English email examples: text: email text. label: 0 for a safe email and 1 for a phishing email. Label Class Examples 0 Safe email 3,000 1 Phishing… See the full description on the dataset page: https://huggingface.co/datasets/jhonrayo99/phishing-email-balanced-6000.texttext-classification1K<n<10K0 likes79 downloads1mo agoHugging Face07ralipanah /email-politeness-corpus Email Politeness Corpus This dataset accompanies the paper: A Synthetic Request–Reply Email Corpus Annotated with Document-Level Politeness and Sentence-Level Face Acts Roshad Alipanah, Valentin Barriere, and Jorge BaierFindings of the Association for Computational Linguistics: EMNLP 2026 The corpus consists of synthetic request–reply emails jointly annotated at two levels: Sentence level: multi-label Face Act annotations grounded in Brown and Levinson's politeness theory.… See the full description on the dataset page: https://huggingface.co/datasets/ralipanah/email-politeness-corpus.tabulartext-classificationn<1K0 likes69 downloads22d agoHugging Face08mindweave /email-campaigns Email Marketing Campaign Analytics Dataset (Free Sample) This is a free sample with 4,025 rows. The full dataset has 56,462 rows across 4 tables. Email campaign performance data for a simulated B2B SaaS company running 120 campaigns over 18 months. 15,000 subscribers across 5 segments, 40,000 email events (sends, opens, clicks, bounces, unsubscribes). Features realistic engagement curves: declining open rates over time, segment-specific behavior, A/B test results, and two… See the full description on the dataset page: https://huggingface.co/datasets/mindweave/email-campaigns.tabulartabular-classification1K<n<10K0 likes64 downloads6mo agoHugging Face09NotShrirang /email-spam-filtertabulartext-classification1K<n<10K11 likes55 downloads1y agoHugging Face10stackscan /email-authentication DMARC and SPF Adoption Among Large Organizations Overview This dataset records which of 36,120 large organizations publish SPF and DMARC records on their primary domain, with firmographic context for each: industry, employee band, country, locality and founding year. SPF lists the servers allowed to send mail for a domain. DMARC tells receiving servers what to do with mail that fails that check, and where to send reports. A domain with SPF but no DMARC has… See the full description on the dataset page: https://huggingface.co/datasets/stackscan/email-authentication.tabulartabular-classification10K<n<100K1 likes45 downloads1mo agoHugging Face11cheekymachine /enron_labeled_emails_with_subjects-llama2-7b_finetuningtexttext-classification1K<n<10K5 likes44 downloads3y agoHugging Face12vaibhavalakshmiravideshik /emapa-uberon-4k Uberon-EMAPA Entity Alignment 4K Uberon-EMAPA Entity Alignment 4K is a biomedical entity alignment benchmark for heterogeneous knowledge graphs. It aligns anatomical entities between the Uber-anatomy Ontology (Uberon) and the Edinburgh Mouse Atlas Project Anatomy ontology (EMAPA) using curated cross-references embedded directly in Uberon. The benchmark is designed for realistic entity alignment research rather than simplified label matching. The gold alignment is embedded in… See the full description on the dataset page: https://huggingface.co/datasets/vaibhavalakshmiravideshik/emapa-uberon-4k.textother10K<n<100K0 likes43 downloads4mo agoHugging Face13Febriyansyah /phishing-emails-multilingual Phishing Emails Multilingual (ID/EN) — Synthetic Dataset sintetis & edukatif 600 email dwibahasa Indonesia 🇮🇩 & English 🇺🇸 untuk riset deteksi phishing — oleh Febriyansyah. ⚠️ Synthetic & edu-defense-only — dibuat untuk pembelajaran defensive security, bukan untuk kampanye nyata. Jangan gunakan untuk aktivitas ilegal. Ringkasan 600 baris — 300 phishing / 300 benign (seimbang), 321 EN / 279 ID Kolom: id (int), language (id/en), text (string, badan email)… See the full description on the dataset page: https://huggingface.co/datasets/Febriyansyah/phishing-emails-multilingual.tabulartext-classificationn<1K0 likes42 downloads16d agoHugging Face14MacLeanLuke /fake-email-campaigntext10K<n<100K1 likes37 downloads3y agoHugging Face15s-emanuilov /rivers-knowledge-base Rivers Knowledge Base Dataset Full structured knowledge base combining DBpedia extractions with LLM-augmented data for U.S. rivers. This repository contains raw DBpedia SPARQL query results, augmented hydrological measurements, alternative river names, and geographic and administrative metadata. The knowledge base serves as the source data for knowledge graph construction used in the Licensing Oracle experiments. Citation @article{ackermann2025stemming, title={Stemming… See the full description on the dataset page: https://huggingface.co/datasets/s-emanuilov/rivers-knowledge-base.text10K<n<100K0 likes37 downloads11mo agoHugging Face16Roy229 /huggingface_filesystem_emails_terminal_6071_sales_orders_1787611667tabularn<1K0 likes37 downloads1mo agoHugging Face17xprilion /email-summary-datasettext10K<n<100K3 likes34 downloads2y agoHugging Face18anilguven /turkish_spam_email Dataset Info Dataset obtained via https://www.kaggle.com/datasets/emrahaydemr/turkish-mail-dataset-normalspam tabulartext-classification1K<n<10K0 likes32 downloads3y agoHugging Face19elenigkove /Email_Intent_Classification Dataset Information This is a dataset of English sentences used in emails with six basic categories: request, informational, transaction, feedback. An example looks as follows: {"Email": "Your subscription renewal is confirmed. Thank you for staying with us!", "Intent": "Transaction"} Dataset Sources Instances generated and annotated by ChatGPT 4. Uses Demo for email intent classification tasks. texttext-classificationn<1K1 likes32 downloads2y agoHugging Face20EmanuelNovelo /guardian_articles_full_contenttext1K<n<10K0 likes29 downloads2y agoHugging Face21EmanuelN /ncdc_lassa_fever_timeseries NCDC Lassa Fever Weekly Timeseries Dataset (Nigeria, 2020–2025) Version: 1.0 Maintainer: Emmanuel Niyi-Oriolowo License: CC BY 4.0 Last Updated: 01-12-2025 1. Overview This repository provides a consolidated and standardized dataset of weekly Lassa fever surveillance data in Nigeria from 2020 to 2025. The dataset is derived from the Nigeria Centre for Disease Control (NCDC) Weekly Epidemiological Reports, which are published as PDF documents. The primary objective of… See the full description on the dataset page: https://huggingface.co/datasets/EmanuelN/ncdc_lassa_fever_timeseries.tabularn<1K4 likes29 downloads10mo agoHugging Face22Dipe00 /Urgency-tone-topic-on-enron_labeled_emails_with_subjects-llama2-7b_finetuningtext1K<n<10K0 likes26 downloads2y agoHugging Face23rtweera /customer_care_emails Dataset Card for customer_care_emails This dataset contains synthetically generated emails that a customer care email unit will receive. Dataset Details Dataset Description This dataset is a synthetically generated dataset using Gemini Pro. It is designed for the following hypothetical scenario. Aetheros is a middleware solutions company for web apps. They have five main services: API development, API Monitoring, IAM, API development language called Mercury… See the full description on the dataset page: https://huggingface.co/datasets/rtweera/customer_care_emails.texttext-classification1K<n<10K8 likes26 downloads2y agoHugging Face24RanjuBiswas /ner-email Overview: This dataset is augmented through llama-3.1:8b. The pourpose is to finetune llm for token classification i.e Email in our case. Following tags are present in dataset: full_name : 1 email : 2 gender : 3 city : 4 country : 5 text1K<n<10K0 likes26 downloads2y agoHugging Face25ClarusC64 /legal-advice-email-risk-option-instruction-coherence-v0.1What this dataset does You receive case position facts used risk analysis options recommendation client instruction consistency flags You decide coherent or incoherent Daily use advice QC risk gap detection instruction capture check contradiction flag tabulartext-classificationn<1K0 likes25 downloads7mo agoHugging Face26W3Genesis /srilankan_email_datasettext1K<n<10K2 likes24 downloads3y agoHugging Face27LIACC /Emakhuwa-loanwords-detection Detecting Loanwords in Emakhuwa Paper: Detecting Loanwords in Emakhuwa: An Extremely Low-Resource {B}antu Language Exhibiting Significant Borrowing from Portuguese @inproceedings{ali-etal-2024-detecting, title = "Detecting Loanwords in Emakhuwa: An Extremely Low-Resource {B}antu Language Exhibiting Significant Borrowing from {P}ortuguese", author = "Ali, Felermino Dario Mario and Lopes Cardoso, Henrique and Sousa-Silva, Rui", booktitle = "Proceedings of… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-loanwords-detection.tabulartext-classification10K<n<100K0 likes24 downloads2y agoHugging Face28ansulev /phishing-email-dataset Phishing Email Dataset This dataset on Hugging Face is a direct copy of the 'Phishing Email Detection' dataset from Kaggle, shared under the GNU Lesser General Public License 3.0. The dataset was originally created by the user 'Cyber Cop' on Kaggle. For complete details, including licensing and usage information, please visit the original Kaggle page. texttext-classification10K<n<100K0 likes24 downloads6mo agoHugging Face29beamstation /restaurant-verified-email-access-in-columbus-ohio-us-173432 Restaurant Verified Email Access in Columbus, Ohio, US Free sample dataset from BeamStation Restaurant Verified Email Access in Columbus, Ohio, US This dataset provides weekly‑verified email addresses for 665 established, independent restaurants (or micro‑chains) located in Columbus, Ohio. Large chain enterprises are excluded, ensuring the list focuses on independent operators. Each record includes a validated email address that has undergone our proprietary verification… See the full description on the dataset page: https://huggingface.co/datasets/beamstation/restaurant-verified-email-access-in-columbus-ohio-us-173432.tabulartabular-classificationn<1K0 likes24 downloads6mo agoHugging Face30Davidglown123 /student_email_priority_dataset_v2tabular1K<n<10K0 likes24 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.