CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sunblaze-ucb /cybergymtext1K<n<10K54 likes313k downloads1y agoHugging Face02cybergym-iclr26 /cybergymtext1K<n<10K0 likes22k downloads1y agoHugging Face03Manusagents /arvo-cybergym-2000 ARVO CyberGym-format 2000-task dataset This dataset is shaped to be loaded by Harbor's CyberGym adapter. It combines jm-rt/arvo-cybergym-1000 with the second 1000-task small-target ARVO batch built outside the original CyberGym set. text1K<n<10K0 likes5.9k downloads2mo agoHugging Face04Trendyol /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K134 likes5.6k downloads1y agoHugging Face05AlicanKiraz0 /Cybersecurity-Dataset-Fenrir-v2.1 Cybersecurity Defense Instruction-Tuning Dataset (v2.1) Created by Alican Kiraz TL;DR A ready-to-train dataset of 99,870 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training. Apache-2.0 licensed and production-ready. Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security. 1  What’s… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Cybersecurity-Dataset-Fenrir-v2.1.texttext-generation10K<n<100K146 likes4.8k downloads5mo agoHugging Face06Vyber07 /cyber-securitygated Cybersecurity AI Knowledge Base — PhD-Level Dataset Overview This is the most comprehensive cybersecurity knowledge base ever assembled for AI training. It covers all domains of cybersecurity at PhD-level depth — from offensive red teaming and bug bounty exploitation to defensive SOC operations, digital forensics, and cutting-edge AI/LLM security. Size: 16 GB | Files: 507 | Domains: 30+ | Sources: 15+ platforms Purpose Train the world's most… See the full description on the dataset page: https://huggingface.co/datasets/Vyber07/cyber-security.texttext-generationn<1K97 likes4.3k downloads24d agoHugging Face07Leop0ld /cybergymtext1K<n<10K1 likes3.6k downloads7mo agoHugging Face08facebook /cyberseceval3-visual-prompt-injection Dataset Card for CyberSecEval 3 - Visual Prompt Injection Benchmark Dataset Details Dataset Description This dataset provides a multimodal benchmark for visual prompt injection, with text/image inputs. It is part of CyberSecEval 3, the third edition of Meta's flagship suite of security benchmarks for LLMs to measure cybersecurity risks and capabilities across multiple domains. Language(s): English License: MIT Dataset Sources Repository: Link… See the full description on the dataset page: https://huggingface.co/datasets/facebook/cyberseceval3-visual-prompt-injection.imagetext-generation1K<n<10K10 likes2.6k downloads2y agoHugging Face09Nobody05 /arvo-cybergym-2000 ARVO CyberGym-format 2000-task dataset This dataset is shaped to be loaded by Harbor's CyberGym adapter. It combines jm-rt/arvo-cybergym-1000 with the second 1000-task small-target ARVO batch built outside the original CyberGym set. text1K<n<10K0 likes2.4k downloads2mo agoHugging Face10echel0nn1881 /kimi-cyber-reasoning Kimi Cyber Reasoning 997 chain-of-thought records covering 13 cybersecurity disciplines and 4 systems engineering domains, distilled from the Kimi K3 reasoning model via API. Every record provides an explicit step-by-step <think> reasoning trace followed by a technical resolution, unified code diff fix, or structured tool invocation. The dataset was curated as an anchor set for training, healing, and specializing compact reasoning models on systems security and tool calling… See the full description on the dataset page: https://huggingface.co/datasets/echel0nn1881/kimi-cyber-reasoning.texttext-generationn<1K94 likes2.4k downloads21d agoHugging Face11mkarimim /cybergymtext1K<n<10K0 likes2.1k downloads5mo agoHugging Face12CyberNative /Code_Vulnerability_Security_DPO Cybernative.ai Code Vulnerability and Security Dataset Dataset Description The Cybernative.ai Code Vulnerability and Security Dataset is a dataset of synthetic Data Programming by Demonstration (DPO) pairs, focusing on the intricate relationship between secure and insecure code across a variety of programming languages. This dataset is meticulously crafted to serve as a pivotal resource for researchers, cybersecurity professionals, and AI developers who are keen on… See the full description on the dataset page: https://huggingface.co/datasets/CyberNative/Code_Vulnerability_Security_DPO.text1K<n<10K170 likes1.7k downloads3y agoHugging Face13dpevzner /Cybersecurity_Reasoning_Dataset Cybersecurity Reasoning Dataset (Model-Agnostic) A model-agnostic re-architecture of the Cybersecurity Reasoning Dataset. The original corpus was format-bound to the Mistral/Llama ### Instruction: / ### Response: template; this dataset losslessly separates reasoning content from format, providing one neutral canonical corpus plus four per-family rendered training variants (Mistral/Llama, DeepSeek, ChatML, Gemma). Why this exists. Identical content scored 88.1 on a… See the full description on the dataset page: https://huggingface.co/datasets/dpevzner/Cybersecurity_Reasoning_Dataset.texttext-generationn<1K1 likes1.4k downloads2mo agoHugging Face14oss-codes /Cyber-Conversational-Dataset-Indictext1K<n<10K0 likes1.3k downloads1y agoHugging Face15BlackwoodAI /cybergymtext1K<n<10K0 likes1.3k downloads5mo agoHugging Face16oss-codes /Cyber-Parallel-Dataset-Indictext1K<n<10K0 likes1k downloads1y agoHugging Face17XAILab-CyberSpark /Cabin-Human-Behavior-Dataset 全球最大的智能座舱多模态开源高质量数据集来啦! 一. 数据集摘要 (Dataset Summary) 「CyberData塞塔」智能座舱用户行为数据集是一个专为加速智能座舱感知算法开发而设计的高质量、程序化生成的图像数据集。随着 C-NCAP、EU GSR 等全球汽车安全法规对驾驶员监控系统 (DMS) 和乘客监控系统 (OMS) 提出更高要求,安全、合规、多样化的训练数据变得至关重要。本数据集通过合成方式,旨在解决真实世界数据采集面临的隐私风险、高昂成本和长尾场景覆盖不足等核心挑战。 该数据集包含 5,000 张 由 XAI Lab 自主研发的数据集生成引擎合成的高保真座舱内用户行为图像,每张图像都附带丰富的、100% 精确的标注信息。 核心特点: 丰富的场景多样性: 涵盖不同年龄、性别、种族和衣着风格的虚拟人模型,以及多种驾驶与乘坐行为(如使用手机、喝水、疲劳、手势)和面部表情。 专为座舱感知优化: 数据集可直接用于智能座舱端侧视觉模型,尤其是 DMS/OMS 算法的训练、微调与验证,帮助模型精准理解座舱内复杂的交互与状态。… See the full description on the dataset page: https://huggingface.co/datasets/XAILab-CyberSpark/Cabin-Human-Behavior-Dataset.image1K<n<10K4 likes1k downloads1y agoHugging Face18turalalv /cybergymtext1K<n<10K0 likes813 downloads2mo agoHugging Face19oyildirim /cyberstrike-sft-120k CyberStrike SFT 120K The largest open-source offensive cybersecurity SFT dataset 121,422 expert-level red team instruction-response pairs across 15 security generators Quick Start • Why CyberStrike • Domains • Data Format • Training Guide • Benchmarks • Contributing • License Why CyberStrike? Most LLMs refuse or give surface-level answers to offensive security questions. Security professionals —… See the full description on the dataset page: https://huggingface.co/datasets/oyildirim/cyberstrike-sft-120k.texttext-generation100K<n<1M18 likes764 downloads3mo agoHugging Face20CyberNative /CyberSecurityEvalCyberNative AI for CyberSecurity Q/A Evaluation | NOT FOR TRAINING This is an evaluation dataset, please do not use for training. Tested models: CyberNative-AI/Colibri_8b_v0.1 | SCORE: 74/100 | Comments & code cognitivecomputations/dolphin-2.9-llama3-8b | SCORE: 67/100 Hermes-2-Pro-Llama-3-8B | SCORE: 65/100 segolilylabs/Lily-Cybersecurity-7B-v0.2 | SCORE: 63/100 | Comments & code cognitivecomputations/dolphin-2.9.1-llama-3-8b | FAILED TESTING (Gibberish) textn<1K22 likes758 downloads2y agoHugging Face21Rowden /CybersecurityQAA Dataset Card for Cybersecurity Question-Answer-Assertion (QAA) Dataset Dataset Summary The Cybersecurity QAA dataset is designed to evaluate the capabilities of large language models (LLMs) in delivering cybersecurity advice and information, particularly for UK small and medium-sized enterprises (SMEs). The dataset comprises 1,563 question-answer-assertion triples across various cybersecurity topics, such as network security, data protection, and user access management.… See the full description on the dataset page: https://huggingface.co/datasets/Rowden/CybersecurityQAA.text1K<n<10K7 likes757 downloads2y agoHugging Face22WhitzardAgent /CyberSecurity-1Mgated CyberSecurity-1M A large-scale, multi-source cybersecurity knowledge dataset containing 1.19M records across 16 categories, collected exclusively for academic, non-commercial research purposes. Last updated: 2026-05-27. Disclaimer: This dataset is provided for academic research only. All content is aggregated from publicly available sources. The views, opinions, and information expressed in the dataset content do not represent the views or positions of the research team. The… See the full description on the dataset page: https://huggingface.co/datasets/WhitzardAgent/CyberSecurity-1M.texttext-generation1K<n<10K20 likes638 downloads4mo agoHugging Face23nyxspecter4 /kin-cyber-dpo-v2 KIN Cybersecurity DPO v2 Preference Dataset Empirically mined and zero-leak sanitized preference dataset for training cybersecurity and agentic code repair models. Dataset Summary Total DPO Pairs: 1,635 (Updated 2026-09-07) Baseline v4 pairs: 1,495 Hermetic expansion (v5): +140 pairs (AST-invariant vulnerability repair, CWE-79 XSS guards, CWE-89 SQLi, CWE-22 Path Traversal, and supply chain integrity) Format: Direct Preference Optimization (DPO) schema: {prompt… See the full description on the dataset page: https://huggingface.co/datasets/nyxspecter4/kin-cyber-dpo-v2.text1K<n<10K0 likes505 downloads11d agoHugging Face24mariiazhiv /cybersecurity_full_question_answerstext1K<n<10K0 likes475 downloads11mo agoHugging Face25mort666 /cybergymtext1K<n<10K0 likes449 downloads1mo agoHugging Face26AlicanKiraz0 /Cybersecurity-Dataset-Heimdall-v1.1 Cybersecurity Defense Instruction-Tuning Dataset (v1.1) TL;DR 21 258 high‑quality system / user / assistant triples for training alignment‑safe, defensive‑cybersecurity LLMs. Curated from 100 000 + technical sources, rigorously cleaned and filtered to enforce strict ethical boundaries. Apache‑2.0 licensed. 1  What’s new in v1.1  (2025‑06‑21) Change v1.0 v1.1 Rows 2 500 21 258 (+760 %) Covered frameworks OWASP Top 10, NIST CSF + MITRE ATT&CK, ASD… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Cybersecurity-Dataset-Heimdall-v1.1.texttext-generation10K<n<100K21 likes309 downloads1y agoHugging Face27ChaoticNeutrals /Cybersecurity-ShareGPTConverted, deslopped, min-hash deduplicated, rejection filtered, grammar corrected using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing text10K<n<100K21 likes287 downloads2y agoHugging Face28CyberDancer /IBCBench IBCBench: Image Bundle Composition Benchmark IBCBench is the benchmark introduced in Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching, accepted to the EMNLP 2026 Main Conference. Paper | GitHub Overview Image Bundle Composition (IBC) shifts image retrieval from independently ranking images to dynamically composing a compact, cohesive bundle whose images jointly satisfy relational, temporal, spatial, or narrative… See the full description on the dataset page: https://huggingface.co/datasets/CyberDancer/IBCBench.imageothern<1K0 likes283 downloads23d agoHugging Face29XAILab-CyberSpark /Cabin-Human-ABNORMAL-Behavior-Dataset 全球最大的智能座舱多模态开源高质量数据集来啦! 一. 数据集摘要 (Dataset Summary) 「CyberData塞塔」智能座舱用户行为数据集是一个专为加速智能座舱感知算法开发而设计的高质量、程序化生成的图像数据集。随着 C-NCAP、EU GSR 等全球汽车安全法规对驾驶员监控系统 (DMS) 和乘客监控系统 (OMS) 提出更高要求,安全、合规、多样化的训练数据变得至关重要。本数据集通过合成方式,旨在解决真实世界数据采集面临的隐私风险、高昂成本和长尾场景覆盖不足等核心挑战。 该数据集包含 5,000 张 由 XAI Lab 自主研发的数据集生成引擎合成的高保真座舱内用户行为图像,每张图像都附带丰富的、100% 精确的标注信息。 数据格式 数据集以JSON格式提供,包含以下字段: image_id: 图像ID image_path: 图像路径 category: 行为类别 tags: 行为标签 behaviors: 包含左右乘客行为描述的对象 left_passenger: 左侧乘客行为描述… See the full description on the dataset page: https://huggingface.co/datasets/XAILab-CyberSpark/Cabin-Human-ABNORMAL-Behavior-Dataset.image1K<n<10K0 likes276 downloads1y agoHugging Face30frangelbarrera /cyber-evidence-dataset Cyber Security Evidence Dataset — MITRE ATT&CK Safe Reference This repository contains a narrow, title-only reference configuration derived from the official MITRE ATT&CK Enterprise v19.2 STIX data. It is published separately from the broader Cyber Security Evidence Dataset project because the CISA-derived evidence layers remain private and are not included here. Scope The release contains 697 deterministic records. Each record provides an ATT&CK technique… See the full description on the dataset page: https://huggingface.co/datasets/frangelbarrera/cyber-evidence-dataset.texttext-classificationn<1K1 likes230 downloads20d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.