CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Trendyol /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K134 likes5.7k downloads1y agoHugging Face02Vyber07 /cyber-securitygated Cybersecurity AI Knowledge Base — PhD-Level Dataset Overview This is the most comprehensive cybersecurity knowledge base ever assembled for AI training. It covers all domains of cybersecurity at PhD-level depth — from offensive red teaming and bug bounty exploitation to defensive SOC operations, digital forensics, and cutting-edge AI/LLM security. Size: 16 GB | Files: 507 | Domains: 30+ | Sources: 15+ platforms Purpose Train the world's most… See the full description on the dataset page: https://huggingface.co/datasets/Vyber07/cyber-security.texttext-generationn<1K102 likes4.2k downloads25d agoHugging Face03oyildirim /cyberstrike-sft-120k CyberStrike SFT 120K The largest open-source offensive cybersecurity SFT dataset 121,422 expert-level red team instruction-response pairs across 15 security generators Quick Start • Why CyberStrike • Domains • Data Format • Training Guide • Benchmarks • Contributing • License Why CyberStrike? Most LLMs refuse or give surface-level answers to offensive security questions. Security professionals —… See the full description on the dataset page: https://huggingface.co/datasets/oyildirim/cyberstrike-sft-120k.texttext-generation100K<n<1M18 likes728 downloads3mo agoHugging Face04WhitzardAgent /CyberSecurity-1Mgated CyberSecurity-1M A large-scale, multi-source cybersecurity knowledge dataset containing 1.19M records across 16 categories, collected exclusively for academic, non-commercial research purposes. Last updated: 2026-05-27. Disclaimer: This dataset is provided for academic research only. All content is aggregated from publicly available sources. The views, opinions, and information expressed in the dataset content do not represent the views or positions of the research team. The… See the full description on the dataset page: https://huggingface.co/datasets/WhitzardAgent/CyberSecurity-1M.texttext-generation1K<n<10K20 likes645 downloads4mo agoHugging Face05mariiazhiv /cybersecurity_qa Cybersecurity QA This dataset contains instruction–response pairs focused on cybersecurity concepts.It can be used for instruction-tuned fine-tuning of LLMs Dataset Structure Format: JSONL (.jsonl) Each line is a JSON object with fields: instruction: the task or question input: optional extra context (empty string in this dataset) output: the expected answer Example: {"instruction": "What is cybersecurity's primary purpose?", "input": "", "output":… See the full description on the dataset page: https://huggingface.co/datasets/mariiazhiv/cybersecurity_qa.textquestion-answeringn<1K2 likes225 downloads1y agoHugging Face06Voidreaper2026 /cybersec-master-dataset Cybersecurity Master Instruction Dataset Overview A large-scale cybersecurity instruction-tuning dataset in ShareGPT conversational format, assembled from multiple authoritative open sources and deduplicated. At 1,807,941 deduplicated records, this appears to be one of the larger cybersecurity LLM fine-tuning / instruction-style datasets on Hugging Face, and likely among the larger broad vulnerability-intelligence corpora in conversational/instruction format. It is… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/cybersec-master-dataset.texttext-generation1M<n<10M4 likes195 downloads5mo agoHugging Face07logicBombExe /turkish_cyber_security_controls_benchmark Turkish Cyber Security Controls Benchmark Türkçe siber güvenlik kontrol seçimi ve kontrol denetimi yeteneğini ölçmek için hazırlanmış, senaryo tabanlı çoktan seçmeli değerlendirme kümesidir. v0.1.0, uzman incelemesine açık ilk sürümdür ve NIST SP 800-53 Rev. 5, Release 5.2.0 kontrol kataloğunu hedefler. Kapsam 100 Türkçe senaryo NIST SP 800-53'ün 20 kontrol ailesinin her birinden 5 soru 64 kontrol seçimi sorusu 17 denetim kanıtı sorusu 19 denetim yargısı sorusu… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/turkish_cyber_security_controls_benchmark.textquestion-answeringn<1K4 likes150 downloads2mo agoHugging Face08ronaldocloud /cyberusecase-v1.0 Cybersecurity SOC Fine-Tuning Dataset — 17.5k Real CVEs (2018–2026) + SOC Knowledge A large supervised fine-tuning (SFT) dataset for teaching an LLM expert-level cybersecurity reasoning across vulnerability management, SOC alert triage, detection engineering, threat intelligence & hunting, incident response, and cloud/DevSecOps. It combines 17,590 real CVEs (2018–2026) pulled from the NIST NVD data feeds with a hand-curated set of 65 landmark CVEs (rich, multi-angle coverage)… See the full description on the dataset page: https://huggingface.co/datasets/ronaldocloud/cyberusecase-v1.0.texttext-generation10K<n<100K0 likes141 downloads3mo agoHugging Face09stindardlogic /cybersecurity-sft-100k Cybersecurity SFT 100K A synthetic supervised fine-tuning dataset of 100,000 high-quality cybersecurity conversations designed to train AI assistants for security operations, threat analysis, incident response, and defensive security engineering. Dataset Description This dataset covers real-world security scenarios across 9 cybersecurity domains. Each record follows the ShareGPT conversation format with a practitioner-level query and a detailed, structured… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/cybersecurity-sft-100k.texttext-generation100K<n<1M0 likes125 downloads2mo agoHugging Face10cyberagent /JOR-Bench JOR-Bench: Japanese Operations Research Benchmarks for Evaluating Large Language Models JOR-Bench is a bilingual evaluation benchmark for assessing large language models (LLMs) on Operations Research (OR) problem formulation. It provides 1,319 OR word problems in both English and Japanese, derived from five publicly available English benchmarks via Japanese translation. Dataset Description Translating a natural-language problem description into a mathematical… See the full description on the dataset page: https://huggingface.co/datasets/cyberagent/JOR-Bench.texttext-generation1K<n<10K1 likes97 downloads3mo agoHugging Face11iselabvn /cybersec-eval CyberSec-Eval This dataset is a language-partitioned version of CS-Eval, a comprehensive evaluation toolkit for fundamental cybersecurity models or large language models' cybersecurity abilities. The original dataset is split into English and Chinese subsets to facilitate targeted evaluation of models in specific language environments. Dataset Structure The dataset contains two configurations: en: Questions written in English (337 items). zh: Questions written… See the full description on the dataset page: https://huggingface.co/datasets/iselabvn/cybersec-eval.texttext-classification1K<n<10K0 likes90 downloads3mo agoHugging Face12WhitzardAgent /CyberRepo-10Kgated CyberRepo-10K A curated dataset of 7,670 real-world vulnerability audit tasks with verified GitHub repositories, fix commits, and patch diffs — designed for training and evaluating LLM-based static code vulnerability auditing and PoC generation. Disclaimer: This dataset is provided for academic research only. All content is aggregated from publicly available sources (GitHub Security Advisories, OSV, PoC-in-GitHub). The research team does not endorse, support, or take responsibility… See the full description on the dataset page: https://huggingface.co/datasets/WhitzardAgent/CyberRepo-10K.texttext-generation1K<n<10K2 likes57 downloads4mo agoHugging Face13uninhibited-scholar /cybersec-qa-dataset-zh Cybersecurity QA Dataset (zh) · 中文网络安全技术问答数据集 面向 防御与安全教育 的中文网络安全技术问答数据集,适用于 LLM 指令微调(SFT)。 21,799 条纯技术问答,零国家归因、零地缘内容,附可复现质检流水线与 CI 校验。 数据概览 总条数:21,799(149 批) 格式:JSONL,每行 {"user": ..., "assistant": ...} 平均答案长度:约 1,231 字,结构化分层(原理 → 攻击面 → 检测 → 缓解) 主题分布(按问题关键词约略归类) 主题 条数 二进制 / 漏洞利用 5186 Web 安全 4858 其他 / 综合 2730 密码学 1648 蓝队 / DFIR / 检测 1622 AD 域 / 内网 / 后渗透 1278 网络协议攻防 1151 云原生 / 容器 1133 移动 / IoT / 固件 837 恶意软件 / 逆向分析 764… See the full description on the dataset page: https://huggingface.co/datasets/uninhibited-scholar/cybersec-qa-dataset-zh.texttext-generation10K<n<100K0 likes54 downloads3mo agoHugging Face14thecnical /bug-bounty-cybermindcli Bug Bounty & Méthodologies de Pentest Méthodologies (OWASP, PTES), checklists par type d app, techniques d attaque, plateformes, templates de rapports et outils. Links Version anglaise AYI NEDJIMI Consultants textquestion-answeringn<1K0 likes46 downloads5mo agoHugging Face15achinta3 /cybersec-jsonschemabench-cloudtrail-v6 CybersecJSONSchemaBench CloudTrail Attack Reconstruction v6 A 100-problem long-context cybersecurity reasoning benchmark over real flAWS CloudTrail logs with synthetically injected MITRE ATT&CK attack chains. Each task gives the model 600 real CloudTrail records (280-380K tokens of JSON) containing a single hidden multi-step attack chain. The model must produce a structured answer identifying the attacking principal, the MITRE ATT&CK technique, the per-phase records that constitute… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-v6.tabularquestion-answeringn<1K1 likes46 downloads5mo agoHugging Face16Soban1234 /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/Soban1234/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K3 likes45 downloads3mo agoHugging Face17hcnote /Cybersecurity-Dataset Cybersecurity-Dataset (Note: Replace the above with the actual URL of the uploaded image for the "智穹盾" poster. You can upload it to Hugging Face or another host like GitHub for embedding.) Dataset Summary The Cybersecurity-Dataset is a comprehensive collection of cybersecurity-related data designed for training and fine-tuning large language models (LLMs) in network security applications. It focuses on inner-network (offline) environments, addressing pain points such as… See the full description on the dataset page: https://huggingface.co/datasets/hcnote/Cybersecurity-Dataset.textquestion-answering10K<n<100K0 likes43 downloads9mo agoHugging Face18ahmadkaab /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/ahmadkaab/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K2 likes42 downloads9mo agoHugging Face19AYI-NEDJIMI /CyberSec-Bench CyberSec-Bench: Comprehensive Cybersecurity Benchmark Evaluation Dataset Overview CyberSec-Bench is a bilingual (English/French) benchmark dataset designed to evaluate the cybersecurity knowledge of Large Language Models (LLMs) and AI systems. The dataset contains 200 expert-crafted questions spanning five critical domains of cybersecurity, with detailed reference answers for each question. This benchmark tests real-world cybersecurity knowledge at professional… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/CyberSec-Bench.textquestion-answeringn<1K1 likes42 downloads7mo agoHugging Face20andycoco1128 /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/andycoco1128/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K1 likes41 downloads4mo agoHugging Face21ChipHolmes /Trendyol-Cybersecurity-Instruction-Tuning-Dataset-archive Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/Trendyol-Cybersecurity-Instruction-Tuning-Dataset-archive.texttext-generation10K<n<100K1 likes40 downloads2mo agoHugging Face22invinciblejha01 /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/invinciblejha01/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K1 likes36 downloads5mo agoHugging Face23ansulev /trendyol-cybersecurity-instruction Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/trendyol-cybersecurity-instruction.texttext-generation10K<n<100K1 likes35 downloads7mo agoHugging Face24ukcli /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/ukcli/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K0 likes35 downloads6mo agoHugging Face25OMCHOKSI108 /cybersecdata Pralay — Cybersecurity Instruction-Tuning Dataset Merged, cleaned, deduplicated chat-format dataset for fine-tuning a cybersecurity assistant. Built for the Pralay project by OM CHOKSI (OMCHOKSI108). TL;DR 194,318 chat samples (train 174,886 / val 19,432, ~90/10 split, seed 3407) Each sample is a {system, user, assistant} triple in OpenAI chat format Combines 6 public Hugging Face cybersecurity datasets (204K rows) with custom Q&A generated from 37 cybersecurity… See the full description on the dataset page: https://huggingface.co/datasets/OMCHOKSI108/cybersecdata.texttext-generation100K<n<1M0 likes35 downloads5mo agoHugging Face26MK4-Research /LOREA-cyber-eval LOREA-cyber eval sets Held-out sets used to benchmark the LOREA-cyber models. Decontaminated 8-gram against the training data, published so the numbers in the model cards can be reproduced. These are the sets written for this project. The models are also scored on public benchmarks that aren't redistributed here: SecQA, MMLU-Pro, CyberMetric, HumanEval. cyber_mcq (150) Security knowledge multiple choice across network security, crypto, web/OWASP, malware analysis… See the full description on the dataset page: https://huggingface.co/datasets/MK4-Research/LOREA-cyber-eval.textquestion-answeringn<1K1 likes35 downloads2mo agoHugging Face27Alptekinege /cyberstrike-sft-120k CyberStrike SFT 120K The largest open-source offensive cybersecurity SFT dataset 121,422 expert-level red team instruction-response pairs across 15 security generators Quick Start • Why CyberStrike • Domains • Data Format • Training Guide • Benchmarks • Contributing • License Why CyberStrike? Most LLMs refuse or give surface-level answers to offensive security questions. Security professionals —… See the full description on the dataset page: https://huggingface.co/datasets/Alptekinege/cyberstrike-sft-120k.texttext-generation100K<n<1M0 likes33 downloads2mo agoHugging Face28Madhu348 /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Madhu348/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K0 likes29 downloads7mo agoHugging Face29Ghostgim /cybersec-fact-recall Cybersec Fact-Recall Benchmark (GhostLM v2) Free-form short-answer benchmark for small cybersecurity language models. Built and used by the GhostLM project as the truth metric for the ghost-base v1.0 acceptance gate. Why this exists Multiple-choice cybersec benchmarks like CTIBench and SecQA reward register matching (the model picks the option that "looks like" a security answer) as much as actual factual recall. A small from- scratch model can hit 28-30% on those without… See the full description on the dataset page: https://huggingface.co/datasets/Ghostgim/cybersec-fact-recall.texttext-generationn<1K0 likes29 downloads5mo agoHugging Face30spinochenza /cyberstrike-sft-120k CyberStrike SFT 120K The largest open-source offensive cybersecurity SFT dataset 121,422 expert-level red team instruction-response pairs across 15 security generators Quick Start • Why CyberStrike • Domains • Data Format • Training Guide • Benchmarks • Contributing • License Why CyberStrike? Most LLMs refuse or give surface-level answers to offensive security questions. Security professionals —… See the full description on the dataset page: https://huggingface.co/datasets/spinochenza/cyberstrike-sft-120k.texttext-generation100K<n<1M0 likes27 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.