CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LingoIITGN /PHINCAbstract Code-mixing is the phenomenon of using more than one language in a sentence. In the multilingual communities, it is a very frequently observed pattern of communication on social media platforms. Flexibility to use multiple languages in one text message might help to communicate efficiently with the target audience. But, the noisy user-generated code-mixed text adds to the challenge of processing and understanding natural language to a much larger extent. Machine translation from… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/PHINC.texttranslation10K<n<100K1 likes3.6k downloads2y agoHugging Face02zefang-liu /phishing-email-dataset Phishing Email Dataset This dataset on Hugging Face is a direct copy of the 'Phishing Email Detection' dataset from Kaggle, shared under the GNU Lesser General Public License 3.0. The dataset was originally created by the user 'Cyber Cop' on Kaggle. For complete details, including licensing and usage information, please visit the original Kaggle page. texttext-classification10K<n<100K38 likes1.1k downloads3y agoHugging Face03philschmid /AIME_1983_2024Disclaimer: This is a Benchmark dataset! Do not using in training! This is the Benchmark of AIME from year 1983~2023, and 2024(part 2). Original: https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions 2024(part 1) can be find at https://huggingface.co/datasets/AI-MO/aimo-validation-aime. tabularn<1K0 likes754 downloads2y agoHugging Face04alexkstern /phishing_urlstext100K<n<1M4 likes593 downloads3y agoHugging Face05Mitake /PhishingURLsANDBenignURLstext100K<n<1M3 likes580 downloads4y agoHugging Face06imanoop7 /phishing_url_classification Phishing URL Classification Dataset This dataset contains URLs labeled as 'Safe' (0) or 'Not Safe' (1) for phishing detection tasks. Dataset Summary This dataset contains URLs labeled for phishing detection tasks. It's designed to help train and evaluate models that can identify potentially malicious URLs. Dataset Creation The dataset was synthetically generated using a custom script that creates both legitimate and potentially phishing URLs. This approach… See the full description on the dataset page: https://huggingface.co/datasets/imanoop7/phishing_url_classification.texttext-classification100K<n<1M5 likes291 downloads2y agoHugging Face07yjkim27 /The-Philosophy-Data-Project About dataset The Philosophy Data Project is a corpus and a set of anaylsis based philosophy texts, totaling over 50 texts and 30 authors, made by Kourosh Alizadeh. school: Broad categorization of which school of thought each book belongs to. Sometimes, this classification can be vague or depend on interpretation. Thankfully, texts in this corpus are all distinctive examples of respective school of thought, so at leat here they are reasonable. sentence_spacy and sentence_str:… See the full description on the dataset page: https://huggingface.co/datasets/yjkim27/The-Philosophy-Data-Project.tabular100K<n<1M11 likes289 downloads3y agoHugging Face08AreLit /PhishNChips PhishNChips: A Benchmark for LLM Email-Agent Security PhishNChips is a large-scale benchmark for evaluating how system prompt configurations influence the security behavior of LLM-based email agents. This repository contains the canonical v5.2 release, featuring 2,000 email stimuli and 220,000 adjudicated model evaluations. Dataset Overview The benchmark measures a critical deployment variable: how strongly an LLM's system prompt shapes its phishing detection capabilities… See the full description on the dataset page: https://huggingface.co/datasets/AreLit/PhishNChips.texttext-classification1K<n<10K1 likes260 downloads5mo agoHugging Face09phil329 /OpenVid-1M-mapping Summary This is the extent dataset proposed in the paper "OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation". OpenVid-1M is a high-quality text-to-video dataset designed for research institutions to enhance video quality, featuring high aesthetics, clarity, and resolution. It can be used for direct training or as a quality tuning complement to other video datasets. New Feature: Video-ZIP mapping files now available for efficient video lookup (see Dataset… See the full description on the dataset page: https://huggingface.co/datasets/phil329/OpenVid-1M-mapping.texttext-to-video1M<n<10M0 likes242 downloads2y agoHugging Face10phihung /titanicThe legendary Titanic dataset from this Kaggle competition tabularn<1K9 likes214 downloads4y agoHugging Face11Rishik001 /PhishingURLDatasetstabular10K<n<100K1 likes120 downloads1y agoHugging Face12locuoco /the-biggest-spam-ham-phish-email-dataset-300000 The Biggest Spam Ham Phish Email Dataset (250000+) This dataset is a large-scale, unified, and deduplicated collection of text messages and emails created for spam, ham, and phishing detection. It has been constructed by combining multiple publicly available and open-source datasets into a single standardized format, making it suitable for machine learning, deep learning, and NLP-based projects. The dataset contains approximately unique 250,000+ samples, covering a diverse range of… See the full description on the dataset page: https://huggingface.co/datasets/locuoco/the-biggest-spam-ham-phish-email-dataset-300000.texttext-classification100K<n<1M0 likes117 downloads7mo agoHugging Face13SM-Bello /PHI-CTRL-F16-Fault-Recovery-Telemetry PHI-CTRL F-16 Actuator Fault Recovery Dataset High-Fidelity JSBSim 6-DOF Telemetry for Physics-Hybrid Self-Healing Flight Control Official verification artifacts of the PHI-CTRL (Physics-Hybrid Integrity Control) architecture — a digital-twin-driven, self-healing flight control framework that actively compensates actuator degradation in real time. Author: Mohammed Bello Sani (SM-Bello) Affiliation: Air Force Institute of Technology (AFIT), Kaduna · Penelope Inc. / PHI Lab… See the full description on the dataset page: https://huggingface.co/datasets/SM-Bello/PHI-CTRL-F16-Fault-Recovery-Telemetry.tabulartime-series-forecasting10K<n<100K0 likes113 downloads15d agoHugging Face14huynq3Cyradar /Phishing_Detection_Datasettext100K<n<1M5 likes104 downloads2y agoHugging Face15Dizzzy0x00 /LLMGen-Phishing-Email-Dataset LLM-Generated Phishing Email Dataset Dataset Description This dataset comprises a collection of phishing and legitimate emails generated using Large Language Models (LLMs), specifically DeepSeek for Chinese emails and OpenAI models for English emails. The primary purpose of this dataset is to facilitate research and development in phishing email detection and classification. The dataset is structured with two key columns: content: The full text content of the email.… See the full description on the dataset page: https://huggingface.co/datasets/Dizzzy0x00/LLMGen-Phishing-Email-Dataset.texttext-classification1K<n<10K2 likes103 downloads9mo agoHugging Face16sayhan /strix-philosophy-qa Strix 134k question-answer pairs based on AiresPucrs' stanford-encyclopedia-philosophy dataset. textquestion-answering100K<n<1M29 likes102 downloads3y agoHugging Face17bgspaditya /phishing-datasettext100K<n<1M5 likes91 downloads3y agoHugging Face18datastax /philosopher-quotes450 quotes by 9 philosophers (50 quotes each), labeled with the author and with a variable number of topic tags. The quotes originally come from https://www.kaggle.com/datasets/mertbozkurt5/quotes-by-philosophers (CC BY-NC-SA 4.0). The text of each quote has been cleaned of soft-hyphens (\xad) and other weird characters. The topic labeling has been done with a default HuggingFace zero-shot classifier pipeline with multi_labels. textn<1K9 likes86 downloads3y agoHugging Face19wangyuancheng /discord-phishing-scam-clean Discord Scam / Clean Messages Dataset 📌 Context This dataset contains real-world messages from my Discord server, labeled to support the fine-tuning of BERT/DistilBERT base models for phishing and scam detection. 💡 Inspiration Traditional Discord moderation bots rely on static keyword rules set by server owners, but scammers easily evade these filters by subtly altering spellings, using homoglyphs, and other tricks.To address this, I built an NLP-powered… See the full description on the dataset page: https://huggingface.co/datasets/wangyuancheng/discord-phishing-scam-clean.texttext-classification1K<n<10K2 likes74 downloads1y agoHugging Face20jhonrayo99 /phishing-email-balanced-6000 Balanced Phishing Email Detection Subset This dataset is a derived, randomly sampled subset of Cyber Cop's Phishing Email Detection dataset on Kaggle. The original dataset is distributed under the GNU Lesser General Public License 3.0. Dataset structure The file phishing_email_subset.csv contains 6,000 English email examples: text: email text. label: 0 for a safe email and 1 for a phishing email. Label Class Examples 0 Safe email 3,000 1 Phishing… See the full description on the dataset page: https://huggingface.co/datasets/jhonrayo99/phishing-email-balanced-6000.texttext-classification1K<n<10K0 likes73 downloads1mo agoHugging Face21mfgiguere /erudit-french-philosophy Dataset Card for Dataset Name Dataset Description Dataset Summary This dataset contains all french philosophy that has been published on erudit.org. It has been generated using a Bs4 web parser that you can find in this repo: https://github.com/MFGiguere/french-philosophy-generator. Supported Tasks and Leaderboards This dataset could be useful for this (non-exhaustive) set of tasks: detect if a text is philosophical or not, generate philosophical… See the full description on the dataset page: https://huggingface.co/datasets/mfgiguere/erudit-french-philosophy.tabular100K<n<1M2 likes69 downloads3y agoHugging Face22philosopher-from-god /ChatGPT-Jailbreak-Prompts-rubend18 Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K2 likes69 downloads1y agoHugging Face23Parv-09 /phi_so101_8bin_v1_trim phi_so101_8bin_v1 — opening-pause trim table Companion to BrutalCaesar/phi_so101_8bin_v1. This is not a dataset. It is a 119-row table plus the script that produced it. The original dataset is unmodified and remains authoritative. Applying this table excludes each episode's pre-teleop dead air as a chunk start point, without deleting a single frame from disk. Why Every episode begins with the arm sitting still while the operator has not yet moved the leader.… See the full description on the dataset page: https://huggingface.co/datasets/Parv-09/phi_so101_8bin_v1_trim.tabularn<1K0 likes65 downloads2mo agoHugging Face24nhatnguyet /cung-phi-bat-trach Cung phi và hướng Bát Trạch Kua number and Bat Trach directions 1. Mô tả · Description Cung phi theo năm sinh và giới tính cho khoảng 1900 tới 2099, kèm bốn hướng tốt và bốn hướng cần tránh. Kua number by birth year and sex for 1900 to 2099, with the four favourable and four unfavourable directions. Số dòng · Rows: 400 Phiên bản · Version: 1.1.0 (2026-09-16) Mã hoá · Encoding: UTF-8 không BOM 2. Cấu trúc · Structure Cột · Column Kiểu · Type… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/cung-phi-bat-trach.tabularn<1K0 likes62 downloads2d agoHugging Face25feti-ai /phiusiil-if3070-stei-itb-2024-2025-1 PhiUSIIL Phishing URL Dataset — IF3070 Coursework Split IF3070 Foundations of Artificial Intelligence · STEI ITB · 2024/2025-1 The PhiUSIIL Phishing URL Dataset as it was distributed for the IF3070 Foundations of Artificial Intelligence course at STEI ITB in the 2024/2025-1 semester — resampled, split into a labelled training file and an unlabelled held-out file, and republished here unmodified. This is the coursework distribution, not the upstream dataset.… See the full description on the dataset page: https://huggingface.co/datasets/feti-ai/phiusiil-if3070-stei-itb-2024-2025-1.tabulartabular-classification100K<n<1M1 likes58 downloads1mo agoHugging Face26kjhq /Philippines-Stock-Symbols-and-Metadata Philippines Stock Symbols & Company Metadata This dataset contains stock symbols and basic company metadata for all listed companies in Philippines.It is updated weekly if new changes are there. 📊 Dataset Contents The dataset is provided as a CSV file with the following columns: Column Description name Full company name ticker Stock ticker symbol (e.g., AAPL, MSFT) market The exchange/market where the stock is listed sector The primary business sector… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/Philippines-Stock-Symbols-and-Metadata.textn<1K0 likes57 downloads1y agoHugging Face27gelcloudy /philippine-elections-2025 Philippine Elections 2025 Dataset (COMELEC) This dataset contains Philippine election results from COMELEC, including both local and overseas voting data. Dataset Structure The dataset is provided in two formats: 1. Combined Datasets combined_local.csv — consolidated local election data combined_overseas.csv — consolidated overseas election data These files merge all regions into a single dataset for easier analysis. 2. Per-Region Datasets… See the full description on the dataset page: https://huggingface.co/datasets/gelcloudy/philippine-elections-2025.tabular10M<n<100M1 likes57 downloads5mo agoHugging Face28ddiask /phishingDatasettext1K<n<10K0 likes55 downloads1y agoHugging Face29jmLuis /MediaFrameCorpus-PhilippineFrameCorpus-CombinedThis training and validation dataset is a combination of Media Frame Corpus and Philippine Frame Corpus, labeled using the Policy Issue Frames Codebook. Train-test split of 80-20. Code_frames column contains annotations following the PolicyIssue Frames Codebook (1-15), wherein at least two(2) annotators agree with the label. The text column contains sentences/phrases from online news articles. The label column is the 0th index code_frames used for training. tabulartext-classification10K<n<100K0 likes49 downloads3y agoHugging Face30debasisdwivedy /Dataset_Philosophy_Ethics_Morality Dataset Card for Dataset Name This dataset card aims to provide reasoning abilitites to LLM models for Philosophical questions. Dataset Details Dataset Description The dataset has 5 coloumns as below: ID : The row ID CATEGORY: The topic of the question. It could relate to morality, ethics, Consciousness etc. QUERY: The question which requires the LLM to think logically. REASONING: The reasoning steps for the LLM to reach to a conclusion. ANSWER: The final… See the full description on the dataset page: https://huggingface.co/datasets/debasisdwivedy/Dataset_Philosophy_Ethics_Morality.textn<1K4 likes44 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.