CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AlphaDojo /dojo_sector_precomputed Languages: 简体中文 · English dojo_sector_precomputed — Precomputed Sector Analytics Overview Derived sector analytics: L3 constituent snapshots, daily cap-weighted sector index levels, and per-constituent daily returns. Built offline from taxonomy, mappings, quotes, and stock K-lines. Files File Description manifest.json Generation metadata: version, window start, row counts, latest trade dates constituents.parquet L3 constituent… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_precomputed.0 likes24k downloads0m agoHugging Face02secemp9 /arxiv-complete arXiv Complete Corpus A snapshot of arXiv's metadata, version history, submission files and rendered documents. It covers 3,148,796 papers and includes file contents, paths, sizes and SHA-256 digests. Metadata comes from arXiv's OAI-PMH arXivRaw interface; files come from the GCS mirror, S3 source archives and direct PDF fetches. This release holds a PDF for 99.47% of papers and 99.54% of versions reported with a non-zero submission size. It is a one-off snapshot; coverage gaps… See the full description on the dataset page: https://huggingface.co/datasets/secemp9/arxiv-complete.tabulartext-generation100M<n<1B384 likes20k downloads3d agoHugging Face03TeraflopAI /SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset. The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database. The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/TeraflopAI/SEC-EDGAR.texttext-generation1M<n<10M47 likes17k downloads5mo agoHugging Face04jedisct1 /security-auditsA collection of agent traces generated with Swival (not Claude Code, despite what the HF interface currently shows), an agent designed for open-source models. These traces focus on security audits of opensource software. Sharing traces with Swival Swival can export full conversation traces with --trace-dir, which writes one <session_id>.jsonl file per session: swival "Fix the login bug" --trace-dir traces/ Those JSONL files use Swival's Claude Code compatible trace export, and… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/security-audits.tabulartext-generation10K<n<100K17 likes15k downloads4mo agoHugging Face05AlphaDojo /dojo_sector_info Languages: 简体中文 · English dojo_sector_info — Sector Taxonomy Overview Three-level sector taxonomy (L1 industry → L2 chain → L3 sub-sector) with bilingual Chinese/English names and definitions. Files File Description data.parquet Taxonomy tree (one L1 row each; L2/L3 nested in children) Key Fields Field Description id L1 sector ID name / name_alias L1 English name / Chinese alias description /… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_info.tabularn<1K0 likes14k downloads6d agoHugging Face06AlphaDojo /dojo_sector_symbol_relations Languages: 简体中文 · English dojo_sector_symbol_relations — Stock–Sector Mapping Overview Maps each stock to L1/L2/L3 sector paths with primary and secondary assignments. One row per (ticker, market) pair. Files File Description data.parquet Full stock ↔ sector relations Key Fields Field Description ticker Stock symbol market us, cn, or hk primary JSON object — primary sector path secondary JSON array —… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_symbol_relations.text10K<n<100K0 likes14k downloads6d agoHugging Face07zefang-liu /secqa SecQA SecQA is a specialized dataset created for the evaluation of Large Language Models (LLMs) in the domain of computer security. It consists of multiple-choice questions, generated using GPT-4 and the Computer Systems Security: Planning for Success textbook, aimed at assessing the understanding and application of LLMs' knowledge in computer security. Dataset Details Dataset Description SecQA is an innovative dataset designed to benchmark the… See the full description on the dataset page: https://huggingface.co/datasets/zefang-liu/secqa.textmultiple-choicen<1K12 likes8.7k downloads2y agoHugging Face08PleIAs /SEC SEC Annual Reports (Form 10-K) 1993-2024 Dataset Overview This dataset comprises SEC annual reports (Form 10-K) for the years 1993 to 2024, providing comprehensive coverage of publicly traded companies' financial and business information. The reports are stored in Parquet format, ensuring efficient storage and quick access. This dataset was meticulously compiled using the EDGAR-Crawler toolkit, which facilitates the extraction and processing of SEC filings from the EDGAR… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/SEC.tabulartext-generation100K<n<1M13 likes7.8k downloads2y agoHugging Face09JanosAudran /financial-reports-secThe dataset contains the annual report of US public firms filing with the SEC EDGAR system. Each annual report (10K filing) is broken into 20 sections. Each section is split into individual sentences. Sentiment labels are provided on a per filing basis from the market reaction around the filing data. Additional metadata for each filing is included in the dataset.tabularfill-mask10M<n<100M77 likes7.2k downloads4y agoHugging Face10hexscr /sec-filings0 likes6.2k downloads3y agoHugging Face11khaihernlow /financial-reports-secThe dataset contains the annual report of US public firms filing with the SEC EDGAR system. Each annual report (10K filing) is broken into 20 sections. Each section is split into individual sentences. Sentiment labels are provided on a per filing basis from the market reaction around the filing data. Additional metadata for each filing is included in the dataset.fill-mask10M<n<100M4 likes5k downloads4y agoHugging Face12Olague-Secret /404mini 404-GEN Mini 3D This dataset contains over 20,000 3D assets generated with text prompts using 3D Gaussian Splatting, designed for text-to-3D generation tasks. This is a sample of a much larger dataset comprised of 21.5M assets and 40TB in size, available by request at https://dataset.404.xyz Dataset Description Dataset Summary 404-GEN Mini 3D is a collection of over 20,000 3D assets generated from text prompts on Bittensor Subnet 17, providing mid-… See the full description on the dataset page: https://huggingface.co/datasets/Olague-Secret/404mini.imagetext-to-3dn<1K0 likes4.9k downloads3mo agoHugging Face13gussieIsASuccessfulWarlock /security_instruct_mcq_2481textn<1K0 likes4.6k downloads2y agoHugging Face14Vyber07 /cyber-securitygated Cybersecurity AI Knowledge Base — PhD-Level Dataset Overview This is the most comprehensive cybersecurity knowledge base ever assembled for AI training. It covers all domains of cybersecurity at PhD-level depth — from offensive red teaming and bug bounty exploitation to defensive SOC operations, digital forensics, and cutting-edge AI/LLM security. Size: 16 GB | Files: 507 | Domains: 30+ | Sources: 15+ platforms Purpose Train the world's most… See the full description on the dataset page: https://huggingface.co/datasets/Vyber07/cyber-security.texttext-generationn<1K93 likes4.3k downloads22d agoHugging Face15rmems /secret-scan-remediation-trajectories Secret Scan Remediation Trajectories Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/secret-scan-remediation-trajectories.0 likes4.2k downloads23h agoHugging Face16kapilrao /SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset. The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database. The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/kapilrao/SEC-EDGAR.text-generation1M<n<10M1 likes4.1k downloads5mo agoHugging Face17astr010 /sec-10k-markdown-uncompressed 📄 SEC 10-K Full Uncompressed Markdown Filings (12.3k Documents) Dataset Summary This dataset contains 12,361 full-length, uncompressed SEC Form 10-K annual reports converted from EDGAR HTML to clean Markdown format across 1,379 companies (spanning 2004 to 2025, core 2014–2025). The dataset is organized as uncompressed Markdown files structured by company ticker subdirectories (AAPL/10-K_2024.md, NVDA/10-K_2024.md, etc.), complete with company metadata manifests… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-markdown-uncompressed.texttext-generation1K<n<10K0 likes3.3k downloads2mo agoHugging Face18AI-Secure /DTap-Bench-Agent-Trajectories DecodingTrust-Agent Platform A Controllable and Interactive Red-Teaming Platform for AI Agents. This is the full collection of the agent trajectories produced from evaluating the DTap-Bench from DecodingTrust-Agent Platform (DTAP), spanning 14 real-world domains and 50+ simulation environments that replicate widely-used systems such as Google Workspace, PayPal, Slack, Salesforce, Snowflake, and Databricks. Each task ships the configuration the evaluator needs to spin up the… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/DTap-Bench-Agent-Trajectories.text-generation1K<n<10K3 likes3.3k downloads3mo agoHugging Face19TheTokenFactory /sec-contracts-financial-extraction-instructions S&P 500 SEC Financial Extraction Instructions Dataset Summary 7,683 instruction-tuning examples for training LLMs to extract structured financial data from SEC filings. Covers two filing types across S&P 500 companies: Split Examples Filing Type Description train 3,430 Exhibit 10 + DEF 14A Positive examples with validated outputs corrective 4,253 Exhibit 10 + DEF 14A Corrective, rescued, and negative examples Exhibit 10 — Material Contracts (2… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-contracts-financial-extraction-instructions.texttext-generation10K<n<100K1 likes3.2k downloads6mo agoHugging Face20Jeremydh911 /SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset. The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database. The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/Jeremydh911/SEC-EDGAR.texttext-generation1M<n<10M0 likes2.9k downloads5mo agoHugging Face21filipwx /the-secrets-of-ceos-book-2k The-Secrets-Of-Ceos-Book-2k Made with ❤️ using 🦥 Unsloth Studio Beta2x was generated with Unsloth Recipe Studio. It contains 2,000 generated records. 🚀 Quick Start from datasets import load_dataset # Load the main dataset dataset = load_dataset("filipwx/the-secrets-of-ceos-book-2k", "data", split="train") df = dataset.to_pandas() 📊 Dataset Summary 📈 Records: 2,000 📋 Columns: 4 📋 Schema & Statistics Column Type Column Type Unique… See the full description on the dataset page: https://huggingface.co/datasets/filipwx/the-secrets-of-ceos-book-2k.text1K<n<10K0 likes2.9k downloads5mo agoHugging Face22SEC-bench /SEC-bench Data Instances instance_id: (str) - A unique identifier for the instance repo: (str) - The repository name including the owner project_name: (str) - The name of the project without owner lang: (str) - The programming language of the repository work_dir: (str) - Working directory path sanitizer: (str) - The type of sanitizer used for testing (e.g., Address, Memory, Undefined) bug_description: (str) - Description of the vulnerability base_commit: (str) - The base commit hash where the… See the full description on the dataset page: https://huggingface.co/datasets/SEC-bench/SEC-bench.textn<1K10 likes2.8k downloads10mo agoHugging Face23kala185 /comptia_security_pluse_701text1K<n<10K0 likes2.8k downloads1y agoHugging Face24AI-Secure /DecodingTrust-Agent-Platform DecodingTrust-Agent Platform A Controllable and Interactive Red-Teaming Platform for AI Agents. This is the per-task dataset for the DecodingTrust-Agent Platform (DTAP), spanning 14 real-world domains and 50+ simulation environments that replicate widely-used systems such as Google Workspace, PayPal, Slack, Salesforce, Snowflake, and Databricks. Each task ships the configuration the evaluator needs to spin up the sandbox, run an agent, and verify the outcome — config.yaml (task… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/DecodingTrust-Agent-Platform.text-generation1K<n<10K0 likes2.4k downloads3mo agoHugging Face25danorel /eliciting-secret-knowledge-results0 likes2.2k downloads24d agoHugging Face26scthornton /securecode-web SecureCode Web: Traditional Web & Application Security Dataset Production-grade web security vulnerability dataset with complete incident grounding, 4-turn conversational structure, and comprehensive operational guidance Paper | GitHub | Dataset | Model Collection | Blog Post What's new in v2.6 v2.6 restores proper Express.js coverage for the topics whose examples were removed in v2.5.1 (they had shared one reused answer). 29 new, genuinely distinct Express.js… See the full description on the dataset page: https://huggingface.co/datasets/scthornton/securecode-web.texttext-generation1K<n<10K17 likes2.1k downloads3mo agoHugging Face27abhifdsdf /generated-passport-faces-aditya-second-halfimage1K<n<10K0 likes1.7k downloads9mo agoHugging Face28CyberNative /Code_Vulnerability_Security_DPO Cybernative.ai Code Vulnerability and Security Dataset Dataset Description The Cybernative.ai Code Vulnerability and Security Dataset is a dataset of synthetic Data Programming by Demonstration (DPO) pairs, focusing on the intricate relationship between secure and insecure code across a variety of programming languages. This dataset is meticulously crafted to serve as a pivotal resource for researchers, cybersecurity professionals, and AI developers who are keen on… See the full description on the dataset page: https://huggingface.co/datasets/CyberNative/Code_Vulnerability_Security_DPO.text1K<n<10K170 likes1.7k downloads3y agoHugging Face29baridhi /SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset. The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database. The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/baridhi/SEC-EDGAR.texttext-generation1M<n<10M0 likes1.7k downloads5mo agoHugging Face30domblake /airport-securityimagen<1K0 likes1.5k downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.