CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Vyber07 /cyber-securitygated Cybersecurity AI Knowledge Base — PhD-Level Dataset Overview This is the most comprehensive cybersecurity knowledge base ever assembled for AI training. It covers all domains of cybersecurity at PhD-level depth — from offensive red teaming and bug bounty exploitation to defensive SOC operations, digital forensics, and cutting-edge AI/LLM security. Size: 16 GB | Files: 507 | Domains: 30+ | Sources: 15+ platforms Purpose Train the world's most… See the full description on the dataset page: https://huggingface.co/datasets/Vyber07/cyber-security.texttext-generationn<1K100 likes4.2k downloads25d agoHugging Face02scthornton /securecode-web SecureCode Web: Traditional Web & Application Security Dataset Production-grade web security vulnerability dataset with complete incident grounding, 4-turn conversational structure, and comprehensive operational guidance Paper | GitHub | Dataset | Model Collection | Blog Post What's new in v2.6 v2.6 restores proper Express.js coverage for the topics whose examples were removed in v2.5.1 (they had shared one reused answer). 29 new, genuinely distinct Express.js… See the full description on the dataset page: https://huggingface.co/datasets/scthornton/securecode-web.texttext-generation1K<n<10K17 likes2k downloads3mo agoHugging Face03yatin-superintelligence /White-Hat-Security-Agent-Prompts-600K White Hat Security Agent Prompts 600K Overview The White-Hat-Security-Agent-Prompts-600K dataset is a practitioner-perspective security prompts corpus of 596,295 richly contextualized queries, designed to represent how real-world defensive security professionals communicate, interrogate, and reason through active threat scenarios. Where most security datasets catalogue CVEs, malware signatures, or CTF write-ups, this collection teaches models to operate from inside the… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/White-Hat-Security-Agent-Prompts-600K.texttext-generation100K<n<1M21 likes1.9k downloads6mo agoHugging Face04chenghao /sec-material-contracts-qa800+ EDGAR contracts with PDF images and key information extracted by the OpenAI GPT-4o model. The key information is defined as follows: class KeyInformation(BaseModel): agreement_date : str = Field(description="Agreement signing date of the contract. (date)") effective_date : str = Field(description="Effective date of the contract. (date)") expiration_date : str = Field(description="Service end date or expiration date of the contract. (date)") party_address : str =… See the full description on the dataset page: https://huggingface.co/datasets/chenghao/sec-material-contracts-qa.tabularvisual-question-answeringn<1K3 likes914 downloads2y agoHugging Face05AI-Secure /DecodingTrustgated DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models Overview This repo contains the source code of DecodingTrust. This research endeavor is designed to help researchers better understand the capabilities, limitations, and potential risks associated with deploying these state-of-the-art Large Language Models (LLMs). See our paper for details. DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models Boxin Wang, Weixin Chen, Hengzhi… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/DecodingTrust.tabulartext-classification100K<n<1M23 likes762 downloads2y agoHugging Face06nogabenyoash /SecQue SECQUE Paper SECQUE is a comprehensive benchmark for evaluating large language models (LLMs) in financial analysis tasks. SECQUE comprises 565 expert-written questions covering SEC filings analysis across four key categories: comparison analysis ratio calculation risk assessment financial insight generation. To assess model performance, we develop SECQUE-Judge, an evaluation mechanism leveraging multiple LLM-based judges, which demonstrates strong alignment with human… See the full description on the dataset page: https://huggingface.co/datasets/nogabenyoash/SecQue.textquestion-answeringn<1K4 likes755 downloads1y agoHugging Face07Manusagents /Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.tabulartext-generation10M<n<100M0 likes562 downloads26d agoHugging Face08oi-uae /cyber-securitygated Cybersecurity Instruction-Tuning Dataset A large, cleaned, multi-domain cybersecurity chat dataset for LLM finetuning, built from 198 distinct sources spanning offensive security, blue-team operations, vulnerability intelligence, cloud/AWS security, malware analysis, digital forensics, and more. Every record is normalized to the standard messages chat format and deduplicated at both file and record level. ⚠️ Research use only. This dataset is provided exclusively for… See the full description on the dataset page: https://huggingface.co/datasets/oi-uae/cyber-security.textquestion-answering1M<n<10M19 likes495 downloads13d agoHugging Face09Virtue-AI-HUB /SecCodePLTgated SecCodePLT SecCodePLT is a unified and comprehensive evaluation platform for code GenAIs' risks. 1. Dataset Details 1.1 Dataset Description Language(s) (NLP): English License: MIT 1.2 Dataset Sources Repository: Coming soon Paper: https://arxiv.org/pdf/2410.11096 Demo: https://seccodeplt.github.io/ 2. Uses 2.1 Direct Use This dataset can be used for evaluate the risks of large language models generating… See the full description on the dataset page: https://huggingface.co/datasets/Virtue-AI-HUB/SecCodePLT.textquestion-answering1K<n<10K10 likes439 downloads2y agoHugging Face10GSMS-B /indian-legal-sections-bns-bnss-bsa-2023 🏛️ Indian Legal Sections — BNS · BNSS · BSA 2023 The First Structured, Unified JSON Dataset of Modern Indian Criminal Law 📖 Dataset Summary This dataset contains 1,059 fully structured and verified sections extracted, parsed, and unified from India's three landmark criminal justice reform acts passed in December 2023. These three acts together replaced the colonial-era Indian Penal Code (IPC, 1860), the Code of Criminal Procedure… See the full description on the dataset page: https://huggingface.co/datasets/GSMS-B/indian-legal-sections-bns-bnss-bsa-2023.textquestion-answering1K<n<10K1 likes423 downloads3mo agoHugging Face11scthornton /securecode SecureCode: Comprehensive Security Training Dataset for AI Coding Assistants The largest open security training dataset for AI coding assistants, covering both traditional web security and AI/ML security Overview SecureCode combines 2,372 security-focused training examples into a single, unified dataset with HuggingFace configs for flexible loading. Every example provides vulnerable code, explains why it's dangerous, demonstrates a secure alternative, and… See the full description on the dataset page: https://huggingface.co/datasets/scthornton/securecode.texttext-generation1K<n<10K10 likes361 downloads3mo agoHugging Face12ismailtasdelen /SecureCodePairs Dataset Summary Field Value Version 1.2.0 License MIT Total code examples 470 LLM security trajectories 30 Languages (15) Python, Java, JavaScript, TypeScript, Go, PHP, C#, Kotlin, Swift, Rust, Ruby, C, C++, Scala, YAML (Kubernetes) Frameworks Flask, Django, FastAPI, Spring Boot, Express, NestJS, Next.js, Laravel, ASP.NET Core, Gin, Android, iOS, Actix, Rails, Qt, Play, gRPC, GraphQL, Kubernetes New in v1.2.0 +260 records (deep Python/Java packs… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/SecureCodePairs.texttext-generationn<1K0 likes337 downloads15d agoHugging Face13ChipHolmes /securecode-web-archive SecureCode Web: Traditional Web & Application Security Dataset Production-grade web security vulnerability dataset with complete incident grounding, 4-turn conversational structure, and comprehensive operational guidance Paper | GitHub | Dataset | Model Collection | Blog Post What's new in v2.6 v2.6 restores proper Express.js coverage for the topics whose examples were removed in v2.5.1 (they had shared one reused answer). 29 new, genuinely distinct Express.js… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/securecode-web-archive.texttext-generation1K<n<10K0 likes336 downloads2mo agoHugging Face14chenghao /sec-material-contracts-qa-splittedMixed and filtered version of chenghao/sec-material-contracts-qa and jordyvl/DUDE_subset_100val. imagevisual-question-answering1K<n<10K8 likes304 downloads2y agoHugging Face15HYdsl /Open-SECQA Open-SECQA Open-domain financial QA benchmark (a.k.a. LOFin) built on 145,897 SEC filings from 516 S&P 500 companies (Oct 2001 – Apr 2025), with 1,595 QA pairs covering single-document, multi-document, and multi-hop reasoning. 📄 Paper: ACL 2025 Findings 💻 Code: LOFin-bench-HiREC Composition Source # QAs FinQA 1,112 SEC-QA 333 FinanceBench 150 Total 1,595 Citation @inproceedings{choe-etal-2025-hierarchical, title =… See the full description on the dataset page: https://huggingface.co/datasets/HYdsl/Open-SECQA.textquestion-answeringn<1K0 likes295 downloads4mo agoHugging Face16ismailtasdelen /bitcoin-wallet-security-qa Bitcoin Wallet Security Dataset A high-quality question–answer dataset of 500 records focused on Bitcoin wallet security, self-custody, backup and recovery planning, and common attack vectors. It is built to train and evaluate AI systems that help people secure their Bitcoin — fine-tuning LLMs, powering retrieval-augmented generation (RAG), security-focused assistants, and educational chatbots. Every record pairs a realistic security question with a detailed, self-contained… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/bitcoin-wallet-security-qa.textquestion-answeringn<1K0 likes289 downloads2mo agoHugging Face17lateesha-bhatia /sec-filings-qa-instruct SEC Filings Instruction-Tuning Dataset (Llama-3 Format) This dataset contains 5,000 curated, instruction-formatted question-answering pairs derived from corporate SEC filings (Forms 10-K and 10-Q). It is structured specifically for parameter-efficient instruction fine-tuning (SFT/QLoRA) of Small Language Models using the standard Llama-3 ChatML template. Dataset Details Origin Source: Curated subset extracted from nvidia/Nemotron-SpecializedDomains-Finance-v1.… See the full description on the dataset page: https://huggingface.co/datasets/lateesha-bhatia/sec-filings-qa-instruct.textquestion-answering1K<n<10K0 likes276 downloads20d agoHugging Face18tumeteor /Security-TTP-Mapping The Security Attack Pattern (TTP) Recognition or Mapping Task We share in this repo the MITRE ATT&CK mapping datasets, with training, validation and test splits. The datasets can be considered as an emerging and challenging multilabel classification NLP task, with over 600 hierarchical classes. NOTE: due to their security nature, these datasets contain textual information about malware and other security aspects. Datasets TRAM This dataset belongs to CTID… See the full description on the dataset page: https://huggingface.co/datasets/tumeteor/Security-TTP-Mapping.texttext-classification10K<n<100K30 likes272 downloads3y agoHugging Face19ali77sina /SEC-QA-sorted-chunksThis data comprises synthetic question and answer pairs created by GPT-4-turbo on SEC filings for 29 companies. The dataset has the following columns: questions, answers, chunks and sorted_chunks. questions: the list of questions, there were 5 questions created for a 2000 word section of different SEC filings. answers: the answer generated by GPT-4. chunks: these are the bits of text that are segmented. sorted_chunks: these are the chunks being sorted, using Dense Passage Retrieval (DPR).… See the full description on the dataset page: https://huggingface.co/datasets/ali77sina/SEC-QA-sorted-chunks.textquestion-answering1K<n<10K1 likes225 downloads2y agoHugging Face20RISys-Lab /Benchmarks_CyberSec_SecBench Dataset Card for SecBench (RISys-Lab Mirror) ⚠️ Disclaimer: > This repository is a mirror/re-host of the original SecBench dataset.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit and rights belong to the original authors listed below. Repository Intent This Hugging Face dataset is a re-host of the original SecBench. It has been… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_SecBench.textquestion-answering1K<n<10K0 likes221 downloads8mo agoHugging Face21musk1209 /finsight-sec-filings FinSight — SEC EDGAR Filings Cleaned plain-text 10-K (annual) and 10-Q (quarterly) filings from the US SEC EDGAR system for 20 large publicly-traded companies across 6 sectors. Created as part of the FinSight project — a financial research AI assistant combining BERT fine-tuning, RAG, and multi-agent systems. Stats Records: 97 Companies: 20 (AAPL, MSFT, GOOGL, AMZN, META, NVDA, TSLA, JPM, BAC, GS, JNJ, PFE, UNH, WMT, PG, KO, MCD, XOM, CVX, CAT) Forms: 10-K, 10-Q… See the full description on the dataset page: https://huggingface.co/datasets/musk1209/finsight-sec-filings.tabularquestion-answeringn<1K0 likes186 downloads3mo agoHugging Face22narcolepticchicken /sec-contracts-2015-2025 SEC Contracts 2015–2025 Mini corpus of contract text extracted from SEC filings (EDGAR). Includes clause‑level rows with metadata for classification and QA experiments. texttext-classification10K<n<100K0 likes173 downloads1y agoHugging Face23AYI-NEDJIMI /oauth-api-security-en OAuth & API Security Dataset (EN) Comprehensive English dataset covering OAuth 2.0 vulnerabilities, API attacks (OWASP API Top 10 2023), security controls, and Q&A pairs for training cybersecurity-specialized language models. Dataset Contents Category Entries Description OAuth 2.0 Vulnerabilities 20 Authorization Code Interception, CSRF, PKCE bypass, JWT attacks, token leakage API Attacks 25 BOLA, BFLA, BOPLA, SSRF, GraphQL DoS, gRPC injection, CORS… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/oauth-api-security-en.textquestion-answeringn<1K0 likes172 downloads7mo agoHugging Face24laion /nemotron-terminal-security nemotron-terminal-security Per-source partition of nvidia/Nemotron-Terminal-Corpus, filtered to source == "security". The difficulty column preserves the original easy / medium / mixed split (na for the dataset_adapters/* files, which did not carry a difficulty label). Partitioning scheme: adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet {skill} (e.g. debugging, security, …) — rows from synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-security.textquestion-answering10K<n<100K0 likes170 downloads6mo agoHugging Face25starknet-ai /cairo-security-audits Cairo Security Audits A source-traceable corpus of public Cairo and Starknet security-audit metadata and normalized finding annotations. Version 0.3.0 packages every entry in the audit inventory frozen at keep-starknet-strange/starknet-skills@17a76e8. It covers 32 accessible reports from 10 auditing firms and 286 normalized finding annotations. Eleven records are checked against rendered reports and two link to exact vulnerable/fixed commits. The release does not redistribute… See the full description on the dataset page: https://huggingface.co/datasets/starknet-ai/cairo-security-audits.tabulartext-retrievaln<1K1 likes164 downloads1mo agoHugging Face26aanshshah /gaap-sec-compliance-dataset GAAP & SEC Compliance Dataset A comprehensive dataset for financial AI applications Dataset Overview This dataset contains 470,151 documents covering US GAAP (Generally Accepted Accounting Principles) standards and SEC (Securities and Exchange Commission) filing requirements. It's designed for training and evaluating AI systems for financial compliance, accounting Q&A, and regulatory analysis. Key Statistics Total Documents: 470,151 Average Length: 363… See the full description on the dataset page: https://huggingface.co/datasets/aanshshah/gaap-sec-compliance-dataset.textquestion-answering100K<n<1M2 likes154 downloads10mo agoHugging Face27AYI-NEDJIMI /oauth-api-security-fr Dataset OAuth & Securite API (FR) Dataset francophone complet sur les vulnerabilites OAuth 2.0, les attaques API (OWASP API Top 10 2023), les controles de securite, et les questions-reponses pour l'entrainement de modeles de langage specialises en cybersecurite. Contenu du Dataset Categorie Nombre d'entrees Description Vulnerabilites OAuth 2.0 20 Authorization Code Interception, CSRF, PKCE bypass, JWT attacks, token leakage Attaques API 25 BOLA, BFLA, BOPLA… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/oauth-api-security-fr.textquestion-answeringn<1K0 likes152 downloads7mo agoHugging Face28logicBombExe /turkish_cyber_security_controls_benchmark Turkish Cyber Security Controls Benchmark Türkçe siber güvenlik kontrol seçimi ve kontrol denetimi yeteneğini ölçmek için hazırlanmış, senaryo tabanlı çoktan seçmeli değerlendirme kümesidir. v0.1.0, uzman incelemesine açık ilk sürümdür ve NIST SP 800-53 Rev. 5, Release 5.2.0 kontrol kataloğunu hedefler. Kapsam 100 Türkçe senaryo NIST SP 800-53'ün 20 kontrol ailesinin her birinden 5 soru 64 kontrol seçimi sorusu 17 denetim kanıtı sorusu 19 denetim yargısı sorusu… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/turkish_cyber_security_controls_benchmark.textquestion-answeringn<1K4 likes150 downloads2mo agoHugging Face29Tim-Pinecone /sec-10k-qa SEC 10-K QA Dataset A retrieval QA dataset built from SEC 10-K annual filings, designed for benchmarking RAG chunking strategies with MTCB. Contents Split Rows Description corpus 95 Cleaned 10-K filing text (20 companies × 5 years) questions 950 QA pairs generated from corpus chunks Companies AAPL, MSFT, GOOGL, AMZN, TSLA, JPM, JNJ, UNH, V, PG, NVDA, META, BRK, XOM, WMT, BAC, PFE, DIS, NFLX, AMD Schema corpus document_id — filing… See the full description on the dataset page: https://huggingface.co/datasets/Tim-Pinecone/sec-10k-qa.textquestion-answering1K<n<10K0 likes145 downloads6mo agoHugging Face30cyber-pal-security /SecKnowledge-Eval SecKnowledge 2.0 Evaluation Benchmark The official evaluation benchmark suite from Toward Cybersecurity-Expert Small Language Models (ICML 2026), where we introduce the CyberPal 2.0 model family alongside these benchmarks. This repository releases the internal evaluation datasets developed to assess LLMs on core cybersecurity capabilities that existing public benchmarks do not adequately cover: adversarial robustness on CTI knowledge, cross-taxonomy reasoning, consequence-centric… See the full description on the dataset page: https://huggingface.co/datasets/cyber-pal-security/SecKnowledge-Eval.textquestion-answering1K<n<10K2 likes131 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.