CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jedisct1 /security-auditsA collection of agent traces generated with Swival (not Claude Code, despite what the HF interface currently shows), an agent designed for open-source models. These traces focus on security audits of opensource software. Sharing traces with Swival Swival can export full conversation traces with --trace-dir, which writes one <session_id>.jsonl file per session: swival "Fix the login bug" --trace-dir traces/ Those JSONL files use Swival's Claude Code compatible trace export, and… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/security-audits.tabulartext-generation10K<n<100K17 likes15k downloads4mo agoHugging Face02gussieIsASuccessfulWarlock /security_instruct_mcq_2481textn<1K0 likes4.6k downloads2y agoHugging Face03Vyber07 /cyber-securitygated Cybersecurity AI Knowledge Base — PhD-Level Dataset Overview This is the most comprehensive cybersecurity knowledge base ever assembled for AI training. It covers all domains of cybersecurity at PhD-level depth — from offensive red teaming and bug bounty exploitation to defensive SOC operations, digital forensics, and cutting-edge AI/LLM security. Size: 16 GB | Files: 507 | Domains: 30+ | Sources: 15+ platforms Purpose Train the world's most… See the full description on the dataset page: https://huggingface.co/datasets/Vyber07/cyber-security.texttext-generationn<1K94 likes4.3k downloads22d agoHugging Face04TheTokenFactory /sec-contracts-financial-extraction-instructions S&P 500 SEC Financial Extraction Instructions Dataset Summary 7,683 instruction-tuning examples for training LLMs to extract structured financial data from SEC filings. Covers two filing types across S&P 500 companies: Split Examples Filing Type Description train 3,430 Exhibit 10 + DEF 14A Positive examples with validated outputs corrective 4,253 Exhibit 10 + DEF 14A Corrective, rescued, and negative examples Exhibit 10 — Material Contracts (2… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-contracts-financial-extraction-instructions.texttext-generation10K<n<100K1 likes3.2k downloads6mo agoHugging Face05SEC-bench /SEC-bench Data Instances instance_id: (str) - A unique identifier for the instance repo: (str) - The repository name including the owner project_name: (str) - The name of the project without owner lang: (str) - The programming language of the repository work_dir: (str) - Working directory path sanitizer: (str) - The type of sanitizer used for testing (e.g., Address, Memory, Undefined) bug_description: (str) - Description of the vulnerability base_commit: (str) - The base commit hash where the… See the full description on the dataset page: https://huggingface.co/datasets/SEC-bench/SEC-bench.textn<1K10 likes2.8k downloads10mo agoHugging Face06CyberNative /Code_Vulnerability_Security_DPO Cybernative.ai Code Vulnerability and Security Dataset Dataset Description The Cybernative.ai Code Vulnerability and Security Dataset is a dataset of synthetic Data Programming by Demonstration (DPO) pairs, focusing on the intricate relationship between secure and insecure code across a variety of programming languages. This dataset is meticulously crafted to serve as a pivotal resource for researchers, cybersecurity professionals, and AI developers who are keen on… See the full description on the dataset page: https://huggingface.co/datasets/CyberNative/Code_Vulnerability_Security_DPO.text1K<n<10K170 likes1.7k downloads3y agoHugging Face07bsebench-org /warwick-second-life-dm-2025-raw First-life and second-life battery degradation mode test data BSEBench status: raw_mirror_pending_validation This repository is a raw mirror of the Mendeley Data dataset Test_Data from Sadia Tasnim Mowri, associated with the University of Warwick. The source description states that the dataset was created to study the influence of first-life degradation mode on second-life performance and degradation, with first-life cells brought to around 80% SoH and then evaluated in second-life… See the full description on the dataset page: https://huggingface.co/datasets/bsebench-org/warwick-second-life-dm-2025-raw.imagen<1K0 likes1.4k downloads5mo agoHugging Face08XuanwuAI /SecEval SecEval: A Comprehensive Benchmark for Evaluating Cybersecurity Knowledge of Foundation Models The advent of large language models has ignited a transformative era for the cybersecurity industry. Pioneering applications are being developed, deployed, and utilized in areas such as cybersecurity knowledge QA, vulnerability hunting, and alert investigation. Various researches have indicated that LLMs primarily acquire their knowledge during the pretraining phase, with fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/XuanwuAI/SecEval.text1K<n<10K26 likes1.3k downloads3y agoHugging Face09pAILabs /infosec-security-qatext10K<n<100K12 likes1k downloads2y agoHugging Face10secmlr /davincitext1M<n<10M0 likes897 downloads2mo agoHugging Face11AI-Secure /DecodingTrustgated DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models Overview This repo contains the source code of DecodingTrust. This research endeavor is designed to help researchers better understand the capabilities, limitations, and potential risks associated with deploying these state-of-the-art Large Language Models (LLMs). See our paper for details. DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models Boxin Wang, Weixin Chen, Hengzhi… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/DecodingTrust.tabulartext-classification100K<n<1M23 likes801 downloads2y agoHugging Face12fgaume /affelnet-paris-secteursCe dépôt contient les secteurs entre les collèges et lycées parisiens Affelnet (Affectation des élèves par le Net) est la procédure informatisée utilisée en France pour affecter les élèves de 3ème dans un lycée de secteur pour leur année de Seconde. L'affectation se base sur un score qui prend en compte les résultats scolaires, la sectorisation géographique, le statut de boursier, et des bonus spécifiques comme le bonus IPS (Indice de Positionnement Social). Description des jeux de… See the full description on the dataset page: https://huggingface.co/datasets/fgaume/affelnet-paris-secteurs.tabular10K<n<100K0 likes787 downloads5mo agoHugging Face13galcan /terraform_sec Terraform Security Dataset A comprehensive dataset of 62,406 Terraform projects analyzed for security vulnerabilities using tfsec. This dataset is designed for training Large Language Models (LLMs) to understand, identify, and fix security issues in Terraform infrastructure-as-code. 📊 Dataset Overview Total Examples: 62,406 Terraform projects Secure Projects: 43,575 (69.8%) Insecure Projects: 18,831 (30.2%) Format: JSONL (JSON Lines) Task: Security analysis and… See the full description on the dataset page: https://huggingface.co/datasets/galcan/terraform_sec.texttext-generation100K<n<1M5 likes762 downloads9mo agoHugging Face14secmlr /starcoderdatatext10M<n<100M0 likes604 downloads2mo agoHugging Face15secmlr /pes2o-v3text10M<n<100M0 likes598 downloads2mo agoHugging Face16nogabenyoash /SecQue SECQUE Paper SECQUE is a comprehensive benchmark for evaluating large language models (LLMs) in financial analysis tasks. SECQUE comprises 565 expert-written questions covering SEC filings analysis across four key categories: comparison analysis ratio calculation risk assessment financial insight generation. To assess model performance, we develop SECQUE-Judge, an evaluation mechanism leveraging multiple LLM-based judges, which demonstrates strong alignment with human… See the full description on the dataset page: https://huggingface.co/datasets/nogabenyoash/SecQue.textquestion-answeringn<1K4 likes498 downloads1y agoHugging Face17natnitaract /exams-basic-and-quantum-cryptography-and-security-latex Open Problem Exams: Cryptography and Security (LaTeX) A curated dataset of open-ended exam problems (with solutions) in cryptography and computer security, formatted in LaTeX. The dataset is sourced from university courses at three institutions. Dataset Overview Institution Files Topics Questions Caltech & TU Delft 8 38 145 EPFL 6 19 86 ETH Zurich 1 14 37 MIT 3 33 79 Total 18 104 347 Difficulty Distribution Institution… See the full description on the dataset page: https://huggingface.co/datasets/natnitaract/exams-basic-and-quantum-cryptography-and-security-latex.textn<1K1 likes488 downloads6mo agoHugging Face18KKing23 /secret-loyalty-competition-data Secret-loyalty organisms — training banks and eval batteries The data behind KKing23/secret-loyalty-competition-organisms. Code and full result trail: github.com/kaustubhkislay/secret-loyalty-competition. Why this exists separately from the adapters. The adapters are reproducible from these banks for the price of GPU time. These banks are not reproducible — they were written by an LLM generator, so regenerating gives different data and every published number becomes… See the full description on the dataset page: https://huggingface.co/datasets/KKing23/secret-loyalty-competition-data.texttext-generation1K<n<10K0 likes462 downloads14d agoHugging Face19GSMS-B /indian-legal-sections-bns-bnss-bsa-2023 🏛️ Indian Legal Sections — BNS · BNSS · BSA 2023 The First Structured, Unified JSON Dataset of Modern Indian Criminal Law 📖 Dataset Summary This dataset contains 1,059 fully structured and verified sections extracted, parsed, and unified from India's three landmark criminal justice reform acts passed in December 2023. These three acts together replaced the colonial-era Indian Penal Code (IPC, 1860), the Code of Criminal Procedure… See the full description on the dataset page: https://huggingface.co/datasets/GSMS-B/indian-legal-sections-bns-bnss-bsa-2023.textquestion-answering1K<n<10K1 likes410 downloads3mo agoHugging Face20OpenClaw /clawhub-security-signals ClawHub Security Signals 🦀 ClawHub | 📝 OpenClaw Blog | 🤗 Hugging Face Blog | 📄 Paper | 📄 Pre-Print ClawHub Security Signals is a sanitized, MIT-licensed security-signals dataset for public OpenClaw agent skills. It captures how an agent-skill registry evaluates trust, provenance, bundled code, and scanner evidence at scale. This dataset was presented in the paper ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree. Paper snapshot: this… See the full description on the dataset page: https://huggingface.co/datasets/OpenClaw/clawhub-security-signals.tabulartext-classification10K<n<100K53 likes400 downloads3mo agoHugging Face21AndeXrd /SecurityQuestionstextn<1K0 likes379 downloads2y agoHugging Face22s2e-lab /SecurityEval Dataset Card for SecurityEval This dataset is from the paper titled SecurityEval Dataset: Mining Vulnerability Examples to Evaluate Machine Learning-Based Code Generation Techniques. The project is accepted for The first edition of the International Workshop on Mining Software Repositories Applications for Privacy and Security (MSR4P&S '22). The paper describes the dataset for evaluating machine learning-based code generation output and the application of the dataset to the code… See the full description on the dataset page: https://huggingface.co/datasets/s2e-lab/SecurityEval.textn<1K10 likes344 downloads3y agoHugging Face23SZLHOLDINGS /szl-second-brain-inrepo SZL Second Brain — in-repo lane (public projection) This is DATA, not a model. A retrieval corpus — the in-repo lane of the SZL Second Brain — published as a Hugging Face dataset. It is part of the SZL Forge Family. Publishing it is repo-existence of a corpus only: it trains nothing, evaluates nothing, serves no inference, and upgrades nothing (Λ stays Conjecture-1). What this is A deterministic BM25-ready corpus assembled only from repo-public text of the Alloy… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/szl-second-brain-inrepo.textn<1K0 likes340 downloads2mo agoHugging Face24secmlr /open-web-mathtext1M<n<10M0 likes330 downloads2mo agoHugging Face25ismailtasdelen /SecureCodePairs Dataset Summary Field Value Version 1.2.0 License MIT Total code examples 470 LLM security trajectories 30 Languages (15) Python, Java, JavaScript, TypeScript, Go, PHP, C#, Kotlin, Swift, Rust, Ruby, C, C++, Scala, YAML (Kubernetes) Frameworks Flask, Django, FastAPI, Spring Boot, Express, NestJS, Next.js, Laravel, ASP.NET Core, Gin, Android, iOS, Actix, Rails, Qt, Play, gRPC, GraphQL, Kubernetes New in v1.2.0 +260 records (deep Python/Java packs… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/SecureCodePairs.texttext-generationn<1K0 likes324 downloads13d agoHugging Face26Hodfa71 /pstu-synthetic-secrets PSTU Synthetic Secrets Dataset Synthetic secrets benchmark for evaluating LLM memorization and unlearning, from the paper: Not All Secrets Are Equal: Type-Aware Unlearning for Language Model Secret Removal Hoda Fakhar — ECML PKDD 2026 Dataset Description 175 synthetic secrets across 25 types, each paired with 100 structurally similar decoys for computing the Carlini exposure metric. All data is synthetically generated. No real credentials, PII, or sensitive information… See the full description on the dataset page: https://huggingface.co/datasets/Hodfa71/pstu-synthetic-secrets.texttext-generationn<1K0 likes269 downloads6mo agoHugging Face27HYdsl /Open-SECQA Open-SECQA Open-domain financial QA benchmark (a.k.a. LOFin) built on 145,897 SEC filings from 516 S&P 500 companies (Oct 2001 – Apr 2025), with 1,595 QA pairs covering single-document, multi-document, and multi-hop reasoning. 📄 Paper: ACL 2025 Findings 💻 Code: LOFin-bench-HiREC Composition Source # QAs FinQA 1,112 SEC-QA 333 FinanceBench 150 Total 1,595 Citation @inproceedings{choe-etal-2025-hierarchical, title =… See the full description on the dataset page: https://huggingface.co/datasets/HYdsl/Open-SECQA.textquestion-answeringn<1K0 likes242 downloads4mo agoHugging Face28starknet-ai /cairo-security-audits Cairo Security Audits A source-traceable corpus of public Cairo and Starknet security-audit metadata and normalized finding annotations. Version 0.3.0 packages every entry in the audit inventory frozen at keep-starknet-strange/starknet-skills@17a76e8. It covers 32 accessible reports from 10 auditing firms and 286 normalized finding annotations. Eleven records are checked against rendered reports and two link to exact vulnerable/fixed commits. The release does not redistribute… See the full description on the dataset page: https://huggingface.co/datasets/starknet-ai/cairo-security-audits.tabulartext-retrievaln<1K1 likes240 downloads29d agoHugging Face29AI-Secure /llm-trustworthy-leaderboard-resultstextn<1K0 likes217 downloads2y agoHugging Face30secmlr /fineweb-edutabular10M<n<100M0 likes216 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.