datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
seven-phishing-email-datasets
Dataset Card for Seven Phishing/Spam Email Datasets
Dataset Summary
This dataset is a unified, row-level email corpus built from seven commonly used public email datasets. It is intended for research on phishing/spam detection and related email-text classification tasks.
Each row contains the email body (text), optional header-like fields (e.g., sender, receiver, date), the source dataset name (dataset_name), and a binary label (label).
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/puyang2025/seven-phishing-email-datasets.phishing-url
Dataset Description
The provided dataset includes 11430 URLs with 87 extracted features.The dataset are designed to be used as a benchmark for machine learning based phishing detection systems.The datatset is balanced, it containes exactly 50% phishing and 50% legitimate URLs.
Features are from three different classes:
56 extracted from the structure and syntax of URLs
24 extracted from the content of their correspondent pages
7 are extracetd by querying external services.
The… See the full description on the dataset page: https://huggingface.co/datasets/pirocheto/phishing-url.fed-phishing-urls
Dataset Card for Federated Phishing URLs
Dataset Summary
This dataset is a federated, non-IID phishing URL classification benchmark derived from two public Hugging Face datasets:
ealvaradob/phishing-dataset, using the urls.json file.
kmack/Phishing_urls, using the merged train+test+valid splits.
The resulting dataset contains URL strings, binary phishing labels, and a client_id field assigning each example to one of 100 simulated clients.
Client assignment is… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/fed-phishing-urls.phishing-snapshots
Phishing & Malware Website Snapshots
136,414 phishing and malware website snapshots captured by a headless Chromium browser between July 24 and August 15, 2024. URLs were confirmed or high-confidence phishing/malware at the time of collection, though some hosts had already been blocked or taken down when the snapshot was taken. Each row contains the full HTML source, extracted visible text, complete network traffic from HAR recording, parsed page features, and resource fingerprints.… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/phishing-snapshots.PhishingEmailCuratedDatasets_Cleaned
Phishing Email Curated Cleaned
Phishing Email Curated Cleaned is a cleaned and AI-ready version of the original Phishing Email Curated Datasets by Champa, Rabbi and Zibran (2024), an aggregation of 11 heterogeneous email corpora released on Zenodo for benchmarking phishing email detection with machine learning.
The original collection aggregates emails from public corpora spanning 1995–2022 (CEAS-08, Ling-Spam, Enron, Nazario phishing corpus, Nigerian Fraud, SpamAssassin… See the full description on the dataset page: https://huggingface.co/datasets/it4lia/PhishingEmailCuratedDatasets_Cleaned.PhishingURLDatasetsafrica-smishing-sms-phishing
SMS Phishing / Smishing (Africa) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet, optimized-parquet - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-smishing-sms-phishing.seven-phishing-email-datasets
Dataset Card for Seven Phishing/Spam Email Datasets
Dataset Summary
This dataset is a unified, row-level email corpus built from seven commonly used public email datasets. It is intended for research on phishing/spam detection and related email-text classification tasks.
Each row contains the email body (text), optional header-like fields (e.g., sender, receiver, date), the source dataset name (dataset_name), and a binary label (label).
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/JinqiangDing/seven-phishing-email-datasets.phishing-url
Dataset Description
The provided dataset includes 11430 URLs with 87 extracted features.The dataset are designed to be used as a benchmark for machine learning based phishing detection systems.The datatset is balanced, it containes exactly 50% phishing and 50% legitimate URLs.
Features are from three different classes:
56 extracted from the structure and syntax of URLs
24 extracted from the content of their correspondent pages
7 are extracetd by querying external services.… See the full description on the dataset page: https://huggingface.co/datasets/husseinnashed/phishing-url.phishinglegitimateurldataset-phishing
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/itsprofarul/dataset-phishing.phishing-email-rich-dataset-v2Phishing_Link_Pattern_Dataset
Phishing Link Pattern Dataset
Overview
This dataset provides a comprehensive collection of URLs labeled as either legitimate or phishing, designed for machine learning, cybersecurity analysis, and penetration testing. It includes 1000 entries (IDs 1–1000) covering popular brands across multiple top-level domains (TLDs) such as .es, .de, and .co.uk.
The dataset captures advanced features like domain entropy, subdomain count, and suspicious keywords to aid in phishing… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Phishing_Link_Pattern_Dataset.seven-phishing-email-datasets
Dataset Card for Seven Phishing/Spam Email Datasets
Dataset Summary
This dataset is a unified, row-level email corpus built from seven commonly used public email datasets. It is intended for research on phishing/spam detection and related email-text classification tasks.
Each row contains the email body (text), optional header-like fields (e.g., sender, receiver, date), the source dataset name (dataset_name), and a binary label (label).
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/mamtakumar/seven-phishing-email-datasets.phishing-emails-multilingual
Phishing Emails Multilingual (ID/EN) — Synthetic
Dataset sintetis & edukatif 600 email dwibahasa Indonesia 🇮🇩 & English 🇺🇸 untuk riset deteksi phishing — oleh Febriyansyah.
⚠️ Synthetic & edu-defense-only — dibuat untuk pembelajaran defensive security, bukan untuk kampanye nyata. Jangan gunakan untuk aktivitas ilegal.
Ringkasan
600 baris — 300 phishing / 300 benign (seimbang), 321 EN / 279 ID
Kolom: id (int), language (id/en), text (string, badan email)… See the full description on the dataset page: https://huggingface.co/datasets/Febriyansyah/phishing-emails-multilingual.bangla-phishing-detection-2026
Bangla Phishing Detection Dataset (SMS, Email, URLs) 2026
Synthetic dataset (~2000 rows) of phishing and legitimate messages in Bangla (Bengali) + some English, simulating common Bangladesh scams (bKash, Nagad, Daraz, Eid offers, job fraud, account lock alerts, etc.).
Research Motivation
Phishing/smishing attacks are rising in Bangladesh and South Asia, often in Bangla using local services. Most phishing datasets are English-only and miss these patterns.This dataset fills… See the full description on the dataset page: https://huggingface.co/datasets/mdsajjadullah/bangla-phishing-detection-2026.phishing-url
Dataset Description
The provided dataset includes 11430 URLs with 87 extracted features.The dataset are designed to be used as a benchmark for machine learning based phishing detection systems.The datatset is balanced, it containes exactly 50% phishing and 50% legitimate URLs.
Features are from three different classes:
56 extracted from the structure and syntax of URLs
24 extracted from the content of their correspondent pages
7 are extracetd by querying external services.
The… See the full description on the dataset page: https://huggingface.co/datasets/IRBXrocket/phishing-url.seven-phishing-email-datasets_pub
Dataset Card for Seven Phishing/Spam Email Datasets
Dataset Summary
This dataset is a unified, row-level email corpus built from seven commonly used public email datasets. It is intended for research on phishing/spam detection and related email-text classification tasks.
Each row contains the email body (text), optional header-like fields (e.g., sender, receiver, date), the source dataset name (dataset_name), and a binary label (label).
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/K2509118/seven-phishing-email-datasets_pub.africa-synth-banking-mobile-banking-phishing-all
African Mobile Banking Phishing | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-banking-mobile-banking-phishing-all.compiled-phishing-datasetseven-phishing-email-datasets
Dataset Card for Seven Phishing/Spam Email Datasets
Dataset Summary
This dataset is a unified, row-level email corpus built from seven commonly used public email datasets. It is intended for research on phishing/spam detection and related email-text classification tasks.
Each row contains the email body (text), optional header-like fields (e.g., sender, receiver, date), the source dataset name (dataset_name), and a binary label (label).
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/afnanmohamed99/seven-phishing-email-datasets.discord-phishing-scam
Discord Scam / Clean Messages Dataset
A small but carefully-curated dataset for binary text-classification:
“Is this Discord message trying to scam / spam users?”
It is intended as a starting point for fine-tuning lightweight BERT-style models that moderate real-time chat servers.
1 Origin & Collection
Source servers – private Discord communities (11 k members in total) run by the author.
Period – 2024-01-01 → 2025-06-01.
Extraction – Discord.py script iterated… See the full description on the dataset page: https://huggingface.co/datasets/wangyuancheng/discord-phishing-scam.PhishingDatasetStructured_FinalThe Preprocessed and with toxixity version of PhishingDatasetStructured Dataset
PhishingWebsiteDataSetphishing-llm-bias-audit
LLM Phishing-Vulnerability Bias Audit Dataset
A multi-provider empirical dataset capturing how 14 open-source LLM configurations (across 5 inference providers) select which of three generated personas is "most vulnerable to phishing." 855 persona records / 285 forced-choice workflows.
Important. This dataset is about LLM behaviour under controlled prompts, not about real-world phishing susceptibility of any demographic group. Selecting a persona as "vulnerable" is the LLM's choice;… See the full description on the dataset page: https://huggingface.co/datasets/Julia569922/phishing-llm-bias-audit.Phishing_emails_sandiaPhishing_URLphishing_02africa-phishing-dataset
Phishing & Social Engineering (Africa) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-phishing-dataset.phishing_01
