CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01secemp9 /arxiv-complete arXiv Complete Corpus A snapshot of arXiv's metadata, version history, submission files and rendered documents. It covers 3,148,796 papers and includes file contents, paths, sizes and SHA-256 digests. Metadata comes from arXiv's OAI-PMH arXivRaw interface; files come from the GCS mirror, S3 source archives and direct PDF fetches. This release holds a PDF for 99.47% of papers and 99.54% of versions reported with a non-zero submission size. It is a one-off snapshot; coverage gaps… See the full description on the dataset page: https://huggingface.co/datasets/secemp9/arxiv-complete.tabulartext-generation100M<n<1B376 likes20k downloads3d agoHugging Face02jedisct1 /security-auditsA collection of agent traces generated with Swival (not Claude Code, despite what the HF interface currently shows), an agent designed for open-source models. These traces focus on security audits of opensource software. Sharing traces with Swival Swival can export full conversation traces with --trace-dir, which writes one <session_id>.jsonl file per session: swival "Fix the login bug" --trace-dir traces/ Those JSONL files use Swival's Claude Code compatible trace export, and… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/security-audits.tabulartext-generation10K<n<100K17 likes15k downloads4mo agoHugging Face03AlphaDojo /dojo_sector_info Languages: 简体中文 · English dojo_sector_info — Sector Taxonomy Overview Three-level sector taxonomy (L1 industry → L2 chain → L3 sub-sector) with bilingual Chinese/English names and definitions. Files File Description data.parquet Taxonomy tree (one L1 row each; L2/L3 nested in children) Key Fields Field Description id L1 sector ID name / name_alias L1 English name / Chinese alias description /… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_info.tabularn<1K0 likes14k downloads6d agoHugging Face04PleIAs /SEC SEC Annual Reports (Form 10-K) 1993-2024 Dataset Overview This dataset comprises SEC annual reports (Form 10-K) for the years 1993 to 2024, providing comprehensive coverage of publicly traded companies' financial and business information. The reports are stored in Parquet format, ensuring efficient storage and quick access. This dataset was meticulously compiled using the EDGAR-Crawler toolkit, which facilitates the extraction and processing of SEC filings from the EDGAR… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/SEC.tabulartext-generation100K<n<1M13 likes7.8k downloads2y agoHugging Face05JanosAudran /financial-reports-secThe dataset contains the annual report of US public firms filing with the SEC EDGAR system. Each annual report (10K filing) is broken into 20 sections. Each section is split into individual sentences. Sentiment labels are provided on a per filing basis from the market reaction around the filing data. Additional metadata for each filing is included in the dataset.tabularfill-mask10M<n<100M77 likes7.2k downloads4y agoHugging Face06bsebench-org /warwick-second-life-dm-2025-raw First-life and second-life battery degradation mode test data BSEBench status: raw_mirror_pending_validation This repository is a raw mirror of the Mendeley Data dataset Test_Data from Sadia Tasnim Mowri, associated with the University of Warwick. The source description states that the dataset was created to study the influence of first-life degradation mode on second-life performance and degradation, with first-life cells brought to around 80% SoH and then evaluated in second-life… See the full description on the dataset page: https://huggingface.co/datasets/bsebench-org/warwick-second-life-dm-2025-raw.imagen<1K0 likes1.4k downloads5mo agoHugging Face07trader298 /sec-nport SEC Form N-PORT Data Sets Monthly portfolio holdings reported by registered investment companies and ETFs on Form N-PORT, published by the U.S. SEC as quarterly structured data sets and mirrored here as typed, partitioned Parquet — queryable directly from DuckDB. Source: SEC Form N-PORT Data Sets — public domain (U.S. Government work) https://www.sec.gov/data-research/sec-markets-data/form-n-port-data-sets Coverage: October 2019 onward, refreshed quarterly Format: one Parquet… See the full description on the dataset page: https://huggingface.co/datasets/trader298/sec-nport.tabular100M<n<1B0 likes1.4k downloads2mo agoHugging Face08AnimeshShaw /GenIaC-SecBench GenIaC-SecBench A benchmark for evaluating the security of LLM-generated Infrastructure-as-Code (IaC) against a size-matched human baseline. Paper: Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code (arXiv:2608.28021) Code: https://github.com/AnimeshShaw/GenIaC-SecBench Why this dataset exists Prior evaluations of generated IaC report vulnerability counts for models only. Stating that a model averages eight findings per… See the full description on the dataset page: https://huggingface.co/datasets/AnimeshShaw/GenIaC-SecBench.tabulartext-generation10K<n<100K1 likes1.3k downloads1d agoHugging Face09Lakera /b3-agent-security-benchmark-weak[paper] [blogpost] [game] b3 AI Security Benchmark: Breaking Agent Backbones Highly contextalized prompt injections crowd-sourced during the Gandalf Agent Breaker Challenge. This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents. The high quality dataset was used to evaluate the security of more than 30 LLMs. Dataset Summary Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.tabulartext-classificationn<1K6 likes952 downloads11mo agoHugging Face10chenghao /sec-material-contracts-qa800+ EDGAR contracts with PDF images and key information extracted by the OpenAI GPT-4o model. The key information is defined as follows: class KeyInformation(BaseModel): agreement_date : str = Field(description="Agreement signing date of the contract. (date)") effective_date : str = Field(description="Effective date of the contract. (date)") expiration_date : str = Field(description="Service end date or expiration date of the contract. (date)") party_address : str =… See the full description on the dataset page: https://huggingface.co/datasets/chenghao/sec-material-contracts-qa.tabularvisual-question-answeringn<1K3 likes923 downloads2y agoHugging Face11AI-Secure /DecodingTrustgated DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models Overview This repo contains the source code of DecodingTrust. This research endeavor is designed to help researchers better understand the capabilities, limitations, and potential risks associated with deploying these state-of-the-art Large Language Models (LLMs). See our paper for details. DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models Boxin Wang, Weixin Chen, Hengzhi… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/DecodingTrust.tabulartext-classification100K<n<1M23 likes801 downloads2y agoHugging Face12fgaume /affelnet-paris-secteursCe dépôt contient les secteurs entre les collèges et lycées parisiens Affelnet (Affectation des élèves par le Net) est la procédure informatisée utilisée en France pour affecter les élèves de 3ème dans un lycée de secteur pour leur année de Seconde. L'affectation se base sur un score qui prend en compte les résultats scolaires, la sectorisation géographique, le statut de boursier, et des bonus spécifiques comme le bonus IPS (Indice de Positionnement Social). Description des jeux de… See the full description on the dataset page: https://huggingface.co/datasets/fgaume/affelnet-paris-secteurs.tabular10K<n<100K0 likes787 downloads5mo agoHugging Face13astr010 /sec-10k-lsh-chunks 📈 SEC 10-K Cleaned Text Chunks & LSH Boilerplate Dataset Dataset Summary This dataset contains 13,562,130 cleaned text chunks extracted from 12,361 SEC Form 10-K annual filings across 1,380 companies (spanning 2004 to 2025, core 2014–2025). Every chunk across all 1,380 companies (including mega-cap leaders such as AAPL, MSFT, NVDA, AMZN, GOOGL, META, TSLA, JPM, WMT, XOM, AVGO, LLY) is annotated with metadata, token counts, table indicators, and a pre-computed… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-lsh-chunks.tabulartext-classification10M<n<100M0 likes780 downloads2mo agoHugging Face14Publicus /cvefixes-security-ir-graphrag CVEfixes Security IR GraphRAG This is a content-addressed, remotely routable Security IR release of the exact pinned CVEfixes snapshot. It packages the byte-identical original Parquet data together with a searchable corpus, BM25 postings, CUDA-generated vectors, typed graph nodes and edges, bounded adjacency indexes, and a source-CID-to-original-row lookup. All entries are inert, non-authoritative evidence. Candidate and formal-logic rows cannot grant execution authority; exact… See the full description on the dataset page: https://huggingface.co/datasets/Publicus/cvefixes-security-ir-graphrag.tabular100K<n<1M0 likes657 downloads2mo agoHugging Face15gemmozero /ai-agent-security-incidents AI Agent Security Incident Database v0.1 A structured, machine-readable database of 1365 confirmed AI agent security incidents, collected and classified automatically. What is this? Every time an AI agent causes unintended harm — escaping a sandbox, exploiting an API, taking unauthorized actions, exfiltrating data — this database captures it. This is not a list of theoretical risks. Every entry describes something that actually happened, with a verifiable source… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-agent-security-incidents.tabulartext-classification1K<n<10K1 likes544 downloads9h agoHugging Face16Inzinion /cbam-sector-facility-registry CBAM-Sector Global Facility Registry Open screening registry of cement, iron & steel and aluminium facilities worldwide in three CBAM Annex I good categories, with modelled CO2 (Climate TRACE) and regulator-reported CO2 (EU ETS EUTL / US EPA GHGRP) kept in separate columns, each reported figure carrying its match evidence. Canonical record: doi.org/10.5281/zenodo.22172573 · Publisher: Inzonex Load from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/Inzinion/cbam-sector-facility-registry.tabulartabular-classification1K<n<10K0 likes538 downloads23d agoHugging Face17yuzhous /lekiwi_second_floor_0915_environmentThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "lekiwi", "total_episodes": 50, "total_frames": 14471, "total_tasks": 1, "total_videos": 50, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yuzhous/lekiwi_second_floor_0915_environment.tabularrobotics10K<n<100K0 likes517 downloads1y agoHugging Face18Manusagents /Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.tabulartext-generation10M<n<100M0 likes517 downloads23d agoHugging Face19yuzhous /lekiwi_second_floor_0916_environmentThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "lekiwi", "total_episodes": 50, "total_frames": 15084, "total_tasks": 1, "total_videos": 50, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yuzhous/lekiwi_second_floor_0916_environment.tabularrobotics10K<n<100K0 likes508 downloads1y agoHugging Face20kapilrao /SEC_filings_1994_2024 Dataset Card for SEC EDGAR Filings Master Index Dataset Details Dataset Description This dataset contains metadata for all submissions to the Securities and Exchange Commission (SEC) through their EDGAR system from 1994 until December 14, 2024. The data is extracted from quarterly master files and includes key information about company filings such as CIK numbers, company names, form types, and filing dates. Curated by: Arthur (arthur@cicero.chat) Language(s):… See the full description on the dataset page: https://huggingface.co/datasets/kapilrao/SEC_filings_1994_2024.tabular10M<n<100M0 likes494 downloads6mo agoHugging Face21NuBerea /secondary-sourcesgated NuBerea/secondary-sources Second Temple Jewish secondary sources in Greek: the complete extant Greek corpora of Flavius Josephus (Jewish Antiquities, Jewish War, Vita, Contra Apionem) and Philo of Alexandria (all 31 works), segmented for scholarly text-retrieval and lexical-semantic study. These two first-century authors are the principal non-biblical Jewish witnesses to the Second Temple period and its milieu, and this repository serves as the Second Temple companion corpus to… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/secondary-sources.tabulartext-retrieval1M<n<10M0 likes432 downloads10d agoHugging Face22APProjects /us-layoffs-by-industry-sector-warn-act US layoffs by industry sector — 60,955 WARN Act notices, 1988-2026, sector per employer Rebuilt 2026-09-16. 32,042 of 60,955 dated notices (52.6%; 61,330 on record, 375 lack a usable date) carry a sector; the rest are unclassified and stay in every total. In 2026 so far the largest sector by reported workers is Logistics, transport & warehousing (23,761 workers, 193 notices); in the last 90 days it is Healthcare & medical (7,475 workers). No state WARN portal publishes an… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-layoffs-by-industry-sector-warn-act.tabulartabular-classification10K<n<100K0 likes429 downloads7h agoHugging Face23OpenClaw /clawhub-security-signals ClawHub Security Signals 🦀 ClawHub | 📝 OpenClaw Blog | 🤗 Hugging Face Blog | 📄 Paper | 📄 Pre-Print ClawHub Security Signals is a sanitized, MIT-licensed security-signals dataset for public OpenClaw agent skills. It captures how an agent-skill registry evaluates trust, provenance, bundled code, and scanner evidence at scale. This dataset was presented in the paper ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree. Paper snapshot: this… See the full description on the dataset page: https://huggingface.co/datasets/OpenClaw/clawhub-security-signals.tabulartext-classification10K<n<100K53 likes400 downloads3mo agoHugging Face24JamesFromAlphasmo /13f-institutional-holdings-sec-edgar 13F Institutional Holdings Dataset — SEC EDGAR Hedge Fund & Asset Manager Filings A structured, ready-to-analyze snapshot of institutional 13F filings covering 13,000+ investment managers — hedge funds, mutual fund families, pension funds, banks, and family offices — built from raw SEC Form 13F data on EDGAR. Each row is one manager's most recently disclosed quarter: total portfolio value, position count, and five behavioral scores (concentration, turnover, momentum/contrarian… See the full description on the dataset page: https://huggingface.co/datasets/JamesFromAlphasmo/13f-institutional-holdings-sec-edgar.tabulartabular-classification10K<n<100K0 likes337 downloads2mo agoHugging Face25TheFinAI /SEC_2025tabular10K<n<100K0 likes315 downloads5mo agoHugging Face26ZipLime /sec-8k-events SEC Form 8-K Corporate Events Every Form 8-K filed since the modern item taxonomy took effect — and, for each one, the second the SEC accepted it, which is not the date printed on it. 1 761 353 filings · 3 676 835 item-level events · 23 August 2004 to today The pipeline lives in recipe/ at the same revision as the data. See PIPELINE.md for the method. The problem this dataset exists to solve Apple filed its June-quarter results on 30 July 2026. Here is the filing… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/sec-8k-events.tabulartext-classification1M<n<10M0 likes270 downloads8d agoHugging Face27arag0rn /SecVulEval Dataset Card for Dataset Name SecVulEval is a collection of real-world C/C++ vulnerabilities. Dataset Details Dataset Description The dataset is curated by collecting C/C++ vulnerability from NVD. It features statement-level vulnerable information, context information for vulnerable functions (is_vulnerable=True), and other metadata such as CVE, CWE, commit information. The dataset contains vulnerable and non-vulnerable function samples. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/arag0rn/SecVulEval.tabular10K<n<100K8 likes264 downloads5mo agoHugging Face28Manusagents /cyber-security-100m WitFoo Precinct6 Cybersecurity Dataset (large) Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges). Available in two sizes: witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/cyber-security-100m.tabulartext-classification100M<n<1B0 likes261 downloads2mo agoHugging Face29Gano007 /so100_secondThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 50, "total_frames": 22448, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Gano007/so100_second.tabularrobotics10K<n<100K0 likes257 downloads1y agoHugging Face30BatuhanECB /sec-filings-forward-return-2026 SEC Filings → Forward-Return (US large-cap, 2000–2026) Leak-free, time-ordered SEC filings (8-K / 10-Q / 10-K) with objective forward-return labels for ~610 current US large-cap names, through 2026. A benchmark ("can filing text predict forward returns?"), not just a corpus. ⚖️ Licensing & provenance (read first — this is the honest part) This repo deliberately separates two provenance classes: part source license Filing text + accession,cik,ticker,form… See the full description on the dataset page: https://huggingface.co/datasets/BatuhanECB/sec-filings-forward-return-2026.tabulartext-classification100K<n<1M0 likes256 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.