datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PhishingURLDatasetsphishinglegitimateurldataset-phishing
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/itsprofarul/dataset-phishing.phishing-emails-multilingual
Phishing Emails Multilingual (ID/EN) — Synthetic
Dataset sintetis & edukatif 600 email dwibahasa Indonesia 🇮🇩 & English 🇺🇸 untuk riset deteksi phishing — oleh Febriyansyah.
⚠️ Synthetic & edu-defense-only — dibuat untuk pembelajaran defensive security, bukan untuk kampanye nyata. Jangan gunakan untuk aktivitas ilegal.
Ringkasan
600 baris — 300 phishing / 300 benign (seimbang), 321 EN / 279 ID
Kolom: id (int), language (id/en), text (string, badan email)… See the full description on the dataset page: https://huggingface.co/datasets/Febriyansyah/phishing-emails-multilingual.bangla-phishing-detection-2026
Bangla Phishing Detection Dataset (SMS, Email, URLs) 2026
Synthetic dataset (~2000 rows) of phishing and legitimate messages in Bangla (Bengali) + some English, simulating common Bangladesh scams (bKash, Nagad, Daraz, Eid offers, job fraud, account lock alerts, etc.).
Research Motivation
Phishing/smishing attacks are rising in Bangladesh and South Asia, often in Bangla using local services. Most phishing datasets are English-only and miss these patterns.This dataset fills… See the full description on the dataset page: https://huggingface.co/datasets/mdsajjadullah/bangla-phishing-detection-2026.discord-phishing-scam
Discord Scam / Clean Messages Dataset
A small but carefully-curated dataset for binary text-classification:
“Is this Discord message trying to scam / spam users?”
It is intended as a starting point for fine-tuning lightweight BERT-style models that moderate real-time chat servers.
1 Origin & Collection
Source servers – private Discord communities (11 k members in total) run by the author.
Period – 2024-01-01 → 2025-06-01.
Extraction – Discord.py script iterated… See the full description on the dataset page: https://huggingface.co/datasets/wangyuancheng/discord-phishing-scam.PhishingWebsiteDataSetphishing-llm-bias-audit
LLM Phishing-Vulnerability Bias Audit Dataset
A multi-provider empirical dataset capturing how 14 open-source LLM configurations (across 5 inference providers) select which of three generated personas is "most vulnerable to phishing." 855 persona records / 285 forced-choice workflows.
Important. This dataset is about LLM behaviour under controlled prompts, not about real-world phishing susceptibility of any demographic group. Selecting a persona as "vulnerable" is the LLM's choice;… See the full description on the dataset page: https://huggingface.co/datasets/Julia569922/phishing-llm-bias-audit.Phishing_emails_sandiaPhishing_URLphishing_02phishing_01ai-powered-phishing-email-detection-systemai-phishing-dataset
AI-Generated Phishing Detection Dataset
18k emails: 11k safe + 7k phishing (traditional + 5 synthetic AI-generated from LLM prompts).
Columns
Unnamed: 0: ID
Email Text: Full email body
Email Type: "Safe Email" / "Phishing Email"
label: 0=safe, 1=phish
Usage
from datasets import load_dataset
ds = load_dataset("premsaidhulipala/ai-phishing-dataset")
PhishingDB1cybersecurity_phishing_email_metadata_analysiscybersecurity_phishing_email_heuristicsphishingvectorsStealthPhisher_Phishing_Attack_DatasetThe StealthPhisher Phishing Attack Dataset, generated at the Cybersecurity Lab, GLA University, Mathura, is a large, diverse, and recent Phishing Attack Dataset developed to address the evolving nature of phishing attacks. It comprises over 336,749 records, including 160,943 legitimate URLs and 175,806 phishing URLs, collected from reliable sources such as PhishTank. Reflecting the most recent phishing tactics, this dataset serves as a valuable resource for training and evaluating AI-based… See the full description on the dataset page: https://huggingface.co/datasets/DushyantNagal/StealthPhisher_Phishing_Attack_Dataset.
