datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
spanish_spear_phishingDataset traducido del inglés al español mediante gpt4o mini.
Los mensajes del dataset contienen:
"email_subject": título del correo, no traducido
"sender_name": nombre del emisor, no traducido
"original_email_body": cuerpo del correo original, no traducido
"translated_email_body": cuerpo del correo traducido
El dataset corresponde al dataset de https://github.com/nahmiasd/Prompted-Contextual-Vectors-for-Spear-Phishing-Detection, el cual esta compuesto de:
"enron_ham": mensajes legítimos del… See the full description on the dataset page: https://huggingface.co/datasets/Darito/spanish_spear_phishing.phishing_benign_email_dataset
Phishing and Benign Email Dataset
This dataset contains a curated collection of phishing and legitimate (benign) emails for use in cybersecurity training, phishing detection models, and email classification systems. Each entry is structured with subject, body, intent, technique, target, and classification label.
📁 Dataset Format
The dataset is stored in .jsonl (JSON Lines) format. Each line is a standalone JSON object.
Fields:
Field
Description
id… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/phishing_benign_email_dataset.Phishing_emails_testphishing-email
CEAS-08 Email Phishing Detection Instruction Dataset
This dataset contains instruction-following conversations for email phishing detection, generated from the CEAS-08 email dataset using multiple large language models. It's designed for fine-tuning conversational AI models on cybersecurity tasks.
Dataset Details
Dataset Description
This dataset transforms raw email data into structured instruction-following conversations where an AI security analyst analyzes… See the full description on the dataset page: https://huggingface.co/datasets/luongnv89/phishing-email.Phishing_Link_Pattern_Dataset
Phishing Link Pattern Dataset
Overview
This dataset provides a comprehensive collection of URLs labeled as either legitimate or phishing, designed for machine learning, cybersecurity analysis, and penetration testing. It includes 1000 entries (IDs 1–1000) covering popular brands across multiple top-level domains (TLDs) such as .es, .de, and .co.uk.
The dataset captures advanced features like domain entropy, subdomain count, and suspicious keywords to aid in phishing… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Phishing_Link_Pattern_Dataset.phishing_legitimatephishing-email-soc-agent
Phishing Email SOC Agent Dataset
A knowledge distillation dataset for training SOC (Security Operations Center) agents to detect and analyze phishing emails using tool-calling capabilities.
Dataset Description
This dataset contains 504 examples of email analysis with real tool calls and responses, designed for fine-tuning LLMs to become phishing detection agents. Each example includes:
Email parsing - Extract headers, URLs, IPs, attachments
Threat intelligence lookup -… See the full description on the dataset page: https://huggingface.co/datasets/Ellbendls/phishing-email-soc-agent.generate_phishing_email_finalPhishing_emailsphishing-texts
Phishing Texts Dataset 🎣
Description:
This dataset is a collection of data designed for training text classifiers capable of determining whether a message or email is a phishing attempt or not.
Dataset Information 📨:
The dataset consists of more than 20,000 entries of text messages, which are potential phishing attempts.
Data is structured in two columns:
text: The text of the message or email.
phising: An indicator of whether the message in the text column… See the full description on the dataset page: https://huggingface.co/datasets/David-Egea/phishing-texts.korean-phishing-email
Korean Phishing Email Detection Dataset (Sample Preview)
Note: This is a sample preview (114 samples) of the full dataset (20,000+ samples) to be released in July 2026 as part of the NIPA Open Source AI/SW Development Support Program.
Dataset Description
Split
File
Samples
Description
train
email_train.jsonl
67
English spam/legitimate emails (Enron-based)
test
email_test.jsonl
17
English spam/legitimate test set
korean
korean_phishing_samples.jsonl
30… See the full description on the dataset page: https://huggingface.co/datasets/joohans/korean-phishing-email.Phishing_hybrid_dbdefendable-pain-phishing-pain-receipt-v0.1
Phishing Pain Receipt
"the hooked" — Mr. Defendable
A free pain-receipt dataset from the DefendableOS ecosystem. 2 rows · ready to read · all cited or graded · CC-BY-4.0.
Part of the 100-pack — 100 free pain-receipt datasets dropped from the Defendable Bakery to the open AI-trust community. Different theme per dataset. Same operator voice across all of them.
Tribunal begins before training. No proof, no honey. To the shed.
What's in here
2 pain receipts themed… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-phishing-pain-receipt-v0.1.phishing_benign_email_dataset
Phishing and Benign Email Dataset
This dataset contains a curated collection of phishing and legitimate (benign) emails for use in cybersecurity training, phishing detection models, and email classification systems. Each entry is structured with subject, body, intent, technique, target, and classification label.
📁 Dataset Format
The dataset is stored in .jsonl (JSON Lines) format. Each line is a standalone JSON object.
Fields:
Field
Description
id… See the full description on the dataset page: https://huggingface.co/datasets/Teddyha/phishing_benign_email_dataset.phishing-detection-yesserPhishing_Link_Pattern_Dataset
Phishing Link Pattern Dataset
Overview
This dataset provides a comprehensive collection of URLs labeled as either legitimate or phishing, designed for machine learning, cybersecurity analysis, and penetration testing. It includes 1000 entries (IDs 1–1000) covering popular brands across multiple top-level domains (TLDs) such as .es, .de, and .co.uk.
The dataset captures advanced features like domain entropy, subdomain count, and suspicious keywords to aid in… See the full description on the dataset page: https://huggingface.co/datasets/Shivahoody007/Phishing_Link_Pattern_Dataset.PhishingDatasetgenerate_phishing_email_1300_rowsPhishingDatasetExtendedphishing-datasetPhishing_Email_For_Finetune_LLMphishingPhishingDatasetTestData
