datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PHINCAbstract
Code-mixing is the phenomenon of using more than one language in a sentence. In the multilingual communities, it is a very frequently observed pattern of communication on social media platforms. Flexibility to use multiple languages in one text message might help to communicate efficiently with the target audience. But, the noisy user-generated code-mixed text adds to the challenge of processing and understanding natural language to a much larger extent. Machine translation from… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/PHINC.phishing-email-dataset
Phishing Email Dataset
This dataset on Hugging Face is a direct copy of the 'Phishing Email Detection' dataset from Kaggle, shared under the GNU Lesser General Public License 3.0. The dataset was originally created by the user 'Cyber Cop' on Kaggle. For complete details, including licensing and usage information, please visit the original Kaggle page.
AIME_1983_2024Disclaimer: This is a Benchmark dataset! Do not using in training!
This is the Benchmark of AIME from year 1983~2023, and 2024(part 2).
Original: https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions
2024(part 1) can be find at https://huggingface.co/datasets/AI-MO/aimo-validation-aime.
phishing_urlsPhishingURLsANDBenignURLsThe-Philosophy-Data-Project
About dataset
The Philosophy Data Project is a corpus and a set of anaylsis based philosophy texts, totaling over 50 texts and 30 authors, made by Kourosh Alizadeh.
school: Broad categorization of which school of thought each book belongs to. Sometimes, this classification can be vague or depend on interpretation. Thankfully, texts in this corpus are all distinctive examples of respective school of thought, so at leat here they are reasonable.
sentence_spacy and sentence_str:… See the full description on the dataset page: https://huggingface.co/datasets/yjkim27/The-Philosophy-Data-Project.phishing_url_classification
Phishing URL Classification Dataset
This dataset contains URLs labeled as 'Safe' (0) or 'Not Safe' (1) for phishing detection tasks.
Dataset Summary
This dataset contains URLs labeled for phishing detection tasks. It's designed to help train and evaluate models that can identify potentially malicious URLs.
Dataset Creation
The dataset was synthetically generated using a custom script that creates both legitimate and potentially phishing URLs. This approach… See the full description on the dataset page: https://huggingface.co/datasets/imanoop7/phishing_url_classification.OpenVid-1M-mapping
Summary
This is the extent dataset proposed in the paper "OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation".
OpenVid-1M is a high-quality text-to-video dataset designed for research institutions to enhance video quality, featuring high aesthetics, clarity, and resolution. It can be used for direct training or as a quality tuning complement to other video datasets.
New Feature: Video-ZIP mapping files now available for efficient video lookup (see Dataset… See the full description on the dataset page: https://huggingface.co/datasets/phil329/OpenVid-1M-mapping.titanicThe legendary Titanic dataset from this Kaggle competition
PHI-SPIKE-C172x-Community-Dataset-v1.0
PHI-SPIKE C172X Community Dataset v1.0
Dataset Summary
PHI-SPIKE C172X Community Dataset v1.0 is a simulation-based aerospace Prognostics and Health Management (PHM) dataset and training-artifact release developed from the PHI-SPIKE C172X research campaign.
The release provides:
JSBSim C172X reference telemetry;
benchmark metadata;
training histories;
trained PyTorch model checkpoints;
per-run evaluation metrics; and
five-seed campaign summaries.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/SM-Bello/PHI-SPIKE-C172x-Community-Dataset-v1.0.PhishNChips
PhishNChips: A Benchmark for LLM Email-Agent Security
PhishNChips is a large-scale benchmark for evaluating how system prompt configurations influence the security behavior of LLM-based email agents. This repository contains the canonical v5.2 release, featuring 2,000 email stimuli and 220,000 adjudicated model evaluations.
Dataset Overview
The benchmark measures a critical deployment variable: how strongly an LLM's system prompt shapes its phishing detection capabilities… See the full description on the dataset page: https://huggingface.co/datasets/AreLit/PhishNChips.Phishing_Detection_Datasetthe-biggest-spam-ham-phish-email-dataset-300000
The Biggest Spam Ham Phish Email Dataset (250000+)
This dataset is a large-scale, unified, and deduplicated collection of text messages and emails created for spam, ham, and phishing detection. It has been constructed by combining multiple publicly available and open-source datasets into a single standardized format, making it suitable for machine learning, deep learning, and NLP-based projects.
The dataset contains approximately unique 250,000+ samples, covering a diverse range of… See the full description on the dataset page: https://huggingface.co/datasets/locuoco/the-biggest-spam-ham-phish-email-dataset-300000.PhishingURLDatasetsPHI-CTRL-F16-Fault-Recovery-Telemetry
PHI-CTRL F-16 Actuator Fault Recovery Dataset
High-Fidelity JSBSim 6-DOF Telemetry for Physics-Hybrid Self-Healing Flight Control
Official verification artifacts of the PHI-CTRL (Physics-Hybrid Integrity Control) architecture — a digital-twin-driven, self-healing flight control framework that actively compensates actuator degradation in real time.
Author: Mohammed Bello Sani (SM-Bello)
Affiliation: Air Force Institute of Technology (AFIT), Kaduna · Penelope Inc. / PHI Lab… See the full description on the dataset page: https://huggingface.co/datasets/SM-Bello/PHI-CTRL-F16-Fault-Recovery-Telemetry.LLMGen-Phishing-Email-Dataset
LLM-Generated Phishing Email Dataset
Dataset Description
This dataset comprises a collection of phishing and legitimate emails generated using Large Language Models (LLMs), specifically DeepSeek for Chinese emails and OpenAI models for English emails. The primary purpose of this dataset is to facilitate research and development in phishing email detection and classification.
The dataset is structured with two key columns:
content: The full text content of the email.… See the full description on the dataset page: https://huggingface.co/datasets/Dizzzy0x00/LLMGen-Phishing-Email-Dataset.phishing-datasetphilosopher-quotes450 quotes by 9 philosophers (50 quotes each), labeled with the author and with a variable number of topic tags.
The quotes originally come from https://www.kaggle.com/datasets/mertbozkurt5/quotes-by-philosophers (CC BY-NC-SA 4.0).
The text of each quote has been cleaned of soft-hyphens (\xad) and other weird characters.
The topic labeling has been done with a default HuggingFace zero-shot classifier pipeline with multi_labels.
strix-philosophy-qa
Strix
134k question-answer pairs based on AiresPucrs' stanford-encyclopedia-philosophy dataset.
discord-phishing-scam-clean
Discord Scam / Clean Messages Dataset
📌 Context
This dataset contains real-world messages from my Discord server, labeled to support the fine-tuning of BERT/DistilBERT base models for phishing and scam detection.
💡 Inspiration
Traditional Discord moderation bots rely on static keyword rules set by server owners, but scammers easily evade these filters by subtly altering spellings, using homoglyphs, and other tricks.To address this, I built an NLP-powered… See the full description on the dataset page: https://huggingface.co/datasets/wangyuancheng/discord-phishing-scam-clean.phishing-email-balanced-6000
Balanced Phishing Email Detection Subset
This dataset is a derived, randomly sampled subset of Cyber Cop's
Phishing Email Detection
dataset on Kaggle. The original dataset is distributed under the GNU Lesser
General Public License 3.0.
Dataset structure
The file phishing_email_subset.csv contains 6,000 English email examples:
text: email text.
label: 0 for a safe email and 1 for a phishing email.
Label
Class
Examples
0
Safe email
3,000
1
Phishing… See the full description on the dataset page: https://huggingface.co/datasets/jhonrayo99/phishing-email-balanced-6000.erudit-french-philosophy
Dataset Card for Dataset Name
Dataset Description
Dataset Summary
This dataset contains all french philosophy that has been published on erudit.org. It has been generated using a Bs4 web parser that you can find in this repo: https://github.com/MFGiguere/french-philosophy-generator.
Supported Tasks and Leaderboards
This dataset could be useful for this (non-exhaustive) set of tasks: detect if a text is philosophical or not, generate philosophical… See the full description on the dataset page: https://huggingface.co/datasets/mfgiguere/erudit-french-philosophy.phi_so101_8bin_v1_trim
phi_so101_8bin_v1 — opening-pause trim table
Companion to BrutalCaesar/phi_so101_8bin_v1.
This is not a dataset. It is a 119-row table plus the script that produced it. The original
dataset is unmodified and remains authoritative. Applying this table excludes each episode's
pre-teleop dead air as a chunk start point, without deleting a single frame from disk.
Why
Every episode begins with the arm sitting still while the operator has not yet moved the leader.… See the full description on the dataset page: https://huggingface.co/datasets/Parv-09/phi_so101_8bin_v1_trim.ChatGPT-Jailbreak-Prompts-rubend18
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
phishinglegitimateurlphiusiil-if3070-stei-itb-2024-2025-1
PhiUSIIL Phishing URL Dataset — IF3070 Coursework Split
IF3070 Foundations of Artificial Intelligence · STEI ITB · 2024/2025-1
The PhiUSIIL Phishing URL Dataset as it was distributed for the IF3070 Foundations of
Artificial Intelligence course at STEI ITB in the 2024/2025-1 semester — resampled, split
into a labelled training file and an unlabelled held-out file, and republished here
unmodified.
This is the coursework distribution, not the upstream dataset.… See the full description on the dataset page: https://huggingface.co/datasets/feti-ai/phiusiil-if3070-stei-itb-2024-2025-1.dataset-phishing
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/itsprofarul/dataset-phishing.phishingDatasetphilippine-elections-2025
Philippine Elections 2025 Dataset (COMELEC)
This dataset contains Philippine election results from COMELEC, including both local and overseas voting data.
Dataset Structure
The dataset is provided in two formats:
1. Combined Datasets
combined_local.csv — consolidated local election data
combined_overseas.csv — consolidated overseas election data
These files merge all regions into a single dataset for easier analysis.
2. Per-Region Datasets… See the full description on the dataset page: https://huggingface.co/datasets/gelcloudy/philippine-elections-2025.cung-phi-bat-trach
Cung phi và hướng Bát Trạch
Kua number and Bat Trach directions
1. Mô tả · Description
Cung phi theo năm sinh và giới tính cho khoảng 1900 tới 2099, kèm bốn hướng tốt và bốn hướng cần tránh.
Kua number by birth year and sex for 1900 to 2099, with the four favourable and four unfavourable directions.
Số dòng · Rows: 400
Phiên bản · Version: 1.1.0 (2026-09-16)
Mã hoá · Encoding: UTF-8 không BOM
2. Cấu trúc · Structure
Cột · Column
Kiểu · Type… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/cung-phi-bat-trach.
