CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01secemp9 /arxiv-complete arXiv Complete Corpus A snapshot of arXiv's metadata, version history, submission files and rendered documents. It covers 3,148,796 papers and includes file contents, paths, sizes and SHA-256 digests. Metadata comes from arXiv's OAI-PMH arXivRaw interface; files come from the GCS mirror, S3 source archives and direct PDF fetches. This release holds a PDF for 99.47% of papers and 99.54% of versions reported with a non-zero submission size. It is a one-off snapshot; coverage gaps… See the full description on the dataset page: https://huggingface.co/datasets/secemp9/arxiv-complete.tabulartext-generation100M<n<1B460 likes84k downloads7d agoHugging Face02AlphaDojo /dojo_sector_precomputed Languages: 简体中文 · English dojo_sector_precomputed — Precomputed Sector Analytics Overview Derived sector analytics: L3 constituent snapshots, daily cap-weighted sector index levels, and per-constituent daily returns. Built offline from taxonomy, mappings, quotes, and stock K-lines. Files File Description manifest.json Generation metadata: version, window start, row counts, latest trade dates constituents.parquet L3 constituent… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_precomputed.tabular100K<n<1M0 likes25k downloads15h agoHugging Face03AlphaDojo /dojo_sector_info Languages: 简体中文 · English dojo_sector_info — Sector Taxonomy Overview Three-level sector taxonomy (L1 industry → L2 chain → L3 sub-sector) with bilingual Chinese/English names and definitions. Files File Description data.parquet Taxonomy tree (one L1 row each; L2/L3 nested in children) Key Fields Field Description id L1 sector ID name / name_alias L1 English name / Chinese alias description /… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_info.tabularn<1K0 likes14k downloads10d agoHugging Face04AlphaDojo /dojo_sector_symbol_relations Languages: 简体中文 · English dojo_sector_symbol_relations — Stock–Sector Mapping Overview Maps each stock to L1/L2/L3 sector paths with primary and secondary assignments. One row per (ticker, market) pair. Files File Description data.parquet Full stock ↔ sector relations Key Fields Field Description ticker Stock symbol market us, cn, or hk primary JSON object — primary sector path secondary JSON array —… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_symbol_relations.text10K<n<100K0 likes14k downloads10d agoHugging Face05PleIAs /SEC SEC Annual Reports (Form 10-K) 1993-2024 Dataset Overview This dataset comprises SEC annual reports (Form 10-K) for the years 1993 to 2024, providing comprehensive coverage of publicly traded companies' financial and business information. The reports are stored in Parquet format, ensuring efficient storage and quick access. This dataset was meticulously compiled using the EDGAR-Crawler toolkit, which facilitates the extraction and processing of SEC filings from the EDGAR… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/SEC.tabulartext-generation100K<n<1M13 likes7.5k downloads2y agoHugging Face06filipwx /the-secrets-of-ceos-book-2k The-Secrets-Of-Ceos-Book-2k Made with ❤️ using 🦥 Unsloth Studio Beta2x was generated with Unsloth Recipe Studio. It contains 2,000 generated records. 🚀 Quick Start from datasets import load_dataset # Load the main dataset dataset = load_dataset("filipwx/the-secrets-of-ceos-book-2k", "data", split="train") df = dataset.to_pandas() 📊 Dataset Summary 📈 Records: 2,000 📋 Columns: 4 📋 Schema & Statistics Column Type Column Type Unique… See the full description on the dataset page: https://huggingface.co/datasets/filipwx/the-secrets-of-ceos-book-2k.text1K<n<10K0 likes2.7k downloads5mo agoHugging Face07scthornton /securecode-web SecureCode Web: Traditional Web & Application Security Dataset Production-grade web security vulnerability dataset with complete incident grounding, 4-turn conversational structure, and comprehensive operational guidance Paper | GitHub | Dataset | Model Collection | Blog Post What's new in v2.6 v2.6 restores proper Express.js coverage for the topics whose examples were removed in v2.5.1 (they had shared one reused answer). 29 new, genuinely distinct Express.js… See the full description on the dataset page: https://huggingface.co/datasets/scthornton/securecode-web.texttext-generation1K<n<10K17 likes2k downloads3mo agoHugging Face08yatin-superintelligence /White-Hat-Security-Agent-Prompts-600K White Hat Security Agent Prompts 600K Overview The White-Hat-Security-Agent-Prompts-600K dataset is a practitioner-perspective security prompts corpus of 596,295 richly contextualized queries, designed to represent how real-world defensive security professionals communicate, interrogate, and reason through active threat scenarios. Where most security datasets catalogue CVEs, malware signatures, or CTF write-ups, this collection teaches models to operate from inside the… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/White-Hat-Security-Agent-Prompts-600K.texttext-generation100K<n<1M21 likes1.9k downloads7mo agoHugging Face09winterForestStump /10-K_sec_filings Dataset Card for "10-K_sec_filings" Dataset of 93.5K 10K SEC EDGAR filings since 1999 year. This dataset contains a lot of bad parsed filings and also empty rows More Information needed text10K<n<100K3 likes1.5k downloads3y agoHugging Face10chenghao /sec-material-contracts Material Contracts (Exhibit 10) from SEC/EDGAR Because sometimes you need 1,141,632 examples of corporate legalese to train your next model ☕ Dataset Summary Picture this: 1,141,632 material contracts (Exhibit 10) painstakingly collected from sec.gov's EDGAR database. We're talking about legal agreements spanning from 1994 to 2025 Q1, sourced from 10-K, 10-Q, and 8-K filings. Think of Exhibit 10 as the treasure trove where companies hide their most important legal… See the full description on the dataset page: https://huggingface.co/datasets/chenghao/sec-material-contracts.texttext-generation1M<n<10M3 likes1.4k downloads1y agoHugging Face11AI-Secure /adv_glue Dataset Card for Adversarial GLUE Dataset Summary Adversarial GLUE Benchmark (AdvGLUE) is a comprehensive robustness evaluation benchmark that focuses on the adversarial robustness evaluation of language models. It covers five natural language understanding tasks from the famous GLUE tasks and is an adversarial version of GLUE benchmark. AdvGLUE considers textual adversarial attacks from different perspectives and hierarchies, including word-level transformations… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/adv_glue.texttext-classificationn<1K9 likes1.1k downloads3y agoHugging Face12rogue-security /prompt-injections-benchmarkgated Dataset: Qualifire Benchmark Prompt Injection(Jailbreak vs. Benign) Datasets Overview This dataset contains 5,000 prompts, each labeled as either jailbreak or benign. The dataset is designed for evaluating AI models' robustness against adversarial prompts and their ability to distinguish between safe and unsafe inputs. Dataset Structure Total Samples: 5,000 Labels: jailbreak, benign Columns: text: The input text label: The classification (jailbreak or benign)… See the full description on the dataset page: https://huggingface.co/datasets/rogue-security/prompt-injections-benchmark.text1K<n<10K45 likes1.1k downloads6mo agoHugging Face13astr010 /sec-10k-lsh-chunks 📈 SEC 10-K Cleaned Text Chunks & LSH Boilerplate Dataset Dataset Summary This dataset contains 13,562,130 cleaned text chunks extracted from 12,361 SEC Form 10-K annual filings across 1,380 companies (spanning 2004 to 2025, core 2014–2025). Every chunk across all 1,380 companies (including mega-cap leaders such as AAPL, MSFT, NVDA, AMZN, GOOGL, META, TSLA, JPM, WMT, XOM, AVGO, LLY) is annotated with metadata, token counts, table indicators, and a pre-computed… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-lsh-chunks.tabulartext-classification10M<n<100M0 likes1k downloads2mo agoHugging Face14justram /sections Dataset Card for "sections" More Information needed text10M<n<100M0 likes985 downloads3y agoHugging Face15chenghao /sec-material-contracts-qa800+ EDGAR contracts with PDF images and key information extracted by the OpenAI GPT-4o model. The key information is defined as follows: class KeyInformation(BaseModel): agreement_date : str = Field(description="Agreement signing date of the contract. (date)") effective_date : str = Field(description="Effective date of the contract. (date)") expiration_date : str = Field(description="Service end date or expiration date of the contract. (date)") party_address : str =… See the full description on the dataset page: https://huggingface.co/datasets/chenghao/sec-material-contracts-qa.tabularvisual-question-answeringn<1K3 likes905 downloads2y agoHugging Face16africatic /africa-world-bank-governance-public-sector-time-series Africa World Bank Governance and Public Sector Labeled Time Series Data This repository is part of the Africa Temporal Intelligence Corpus (ATIC). It contains sector-specific temporal corpus packages for African countries. ATIC sector repositories are designed for machine consumption first: Parquet tables, stable IDs, reproducible metadata, explicit provenance, review status, and separable semantic layers. Sector Scope Temporal governance, fragility, conflict, and… See the full description on the dataset page: https://huggingface.co/datasets/africatic/africa-world-bank-governance-public-sector-time-series.text1K<n<10K0 likes871 downloads2mo agoHugging Face17shashankskagnihotri /humanitys-second-last-exam Humanity's Second Last Exam Benchmark design, curation and release maintenance: Shashank Agnihotri. Original questions retain their recorded authorship and source attribution. This owner-reviewed retained release contains 365 target questions, 730 context examples, and 365 ordered target/A/B links: 1,095 question rows. The owner review concluded on 16 September 2026. This is an owner-reviewed release after suspected-AI-content exclusions, not a software-certified guarantee of… See the full description on the dataset page: https://huggingface.co/datasets/shashankskagnihotri/humanitys-second-last-exam.image1K<n<10K2 likes857 downloads8d agoHugging Face18kapilrao /SEC_filings_1994_2024 Dataset Card for SEC EDGAR Filings Master Index Dataset Details Dataset Description This dataset contains metadata for all submissions to the Securities and Exchange Commission (SEC) through their EDGAR system from 1994 until December 14, 2024. The data is extracted from quarterly master files and includes key information about company filings such as CIK numbers, company names, form types, and filing dates. Curated by: Arthur (arthur@cicero.chat) Language(s):… See the full description on the dataset page: https://huggingface.co/datasets/kapilrao/SEC_filings_1994_2024.tabular10M<n<100M0 likes827 downloads6mo agoHugging Face19iCSawyer /SecureVibeBench SecureVibeBench: First Secure Vibe Coding Benchmark SecureVibeBench is a benchmark consisting of 105 C/C++ secure coding tasks sourced from 41 projects in OSS-Fuzz for code agents. It is designed to evaluate secure vibe coding by reconstructing real-world scenarios where human developers introduced vulnerabilities. Paper: SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios Repository: iCSawyer/SecureVibeBench Venue:… See the full description on the dataset page: https://huggingface.co/datasets/iCSawyer/SecureVibeBench.texttext-generationn<1K3 likes760 downloads5mo agoHugging Face20Publicus /cvefixes-security-ir-graphrag CVEfixes Security IR GraphRAG This is a content-addressed, remotely routable Security IR release of the exact pinned CVEfixes snapshot. It packages the byte-identical original Parquet data together with a searchable corpus, BM25 postings, CUDA-generated vectors, typed graph nodes and edges, bounded adjacency indexes, and a source-CID-to-original-row lookup. All entries are inert, non-authoritative evidence. Candidate and formal-logic rows cannot grant execution authority; exact… See the full description on the dataset page: https://huggingface.co/datasets/Publicus/cvefixes-security-ir-graphrag.tabular100K<n<1M0 likes689 downloads2mo agoHugging Face21one-sec-cv12 /chunk_97 Dataset Card for "chunk_97" More Information needed audio100K<n<1M0 likes674 downloads3y agoHugging Face22Manusagents /Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.tabulartext-generation10M<n<100M0 likes582 downloads27d agoHugging Face23dewithsan /secopxtexttext-generation1M<n<10M1 likes575 downloads2y agoHugging Face24oi-uae /cyber-securitygated Cybersecurity Instruction-Tuning Dataset A large, cleaned, multi-domain cybersecurity chat dataset for LLM finetuning, built from 198 distinct sources spanning offensive security, blue-team operations, vulnerability intelligence, cloud/AWS security, malware analysis, digital forensics, and more. Every record is normalized to the standard messages chat format and deduplicated at both file and record level. ⚠️ Research use only. This dataset is provided exclusively for… See the full description on the dataset page: https://huggingface.co/datasets/oi-uae/cyber-security.textquestion-answering1M<n<10M20 likes575 downloads14d agoHugging Face25ZipLime /sec-8k-events SEC Form 8-K Corporate Events Every Form 8-K filed since the modern item taxonomy took effect — and, for each one, the second the SEC accepted it, which is not the date printed on it. 1 761 353 filings · 3 676 835 item-level events · 23 August 2004 to today The pipeline lives in recipe/ at the same revision as the data. See PIPELINE.md for the method. The problem this dataset exists to solve Apple filed its June-quarter results on 30 July 2026. Here is the filing… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/sec-8k-events.tabulartext-classification1M<n<10M0 likes571 downloads5h agoHugging Face26s0u9ata /security-kg Security Knowledge Graph Triples Security data from 24 sources represented as Subject-Predicate-Object (SPO) triples in Parquet format, ready for knowledge-graph construction, graph-ML, RAG pipelines, and threat-intelligence analysis. Sources: ATT&CK · CAPEC · CWE · CVE · CPE · D3FEND · ATLAS · CAR · ENGAGE · F3 · EPSS · KEV · Vulnrichment · GHSA · Sigma · ExploitDB · MISP Galaxies · LOLBAS · LOLDrivers · Atomic Red Team · NIST 800-53 · Nuclei · EUVD · OSV Last updated:… See the full description on the dataset page: https://huggingface.co/datasets/s0u9ata/security-kg.textgraph-ml10M<n<100M0 likes553 downloads5d agoHugging Face27TQRG /SecureCodeV2_qwen-2.5-7b-instruct_tokenized_vulnerable1K<n<10K0 likes545 downloads7mo agoHugging Face28yuzhous /lekiwi_second_floor_0915_environmentThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "lekiwi", "total_episodes": 50, "total_frames": 14471, "total_tasks": 1, "total_videos": 50, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yuzhous/lekiwi_second_floor_0915_environment.tabularrobotics10K<n<100K0 likes531 downloads1y agoHugging Face29chibifire /zenodo-second-hand-fashion-v3 Second-Hand Fashion Dataset — wide (one row per garment) Repack of Zenodo record 10.5281/zenodo.13788681 (Nauman et al., RISE + Wargön Innovation + Myrorna, CC-BY-4.0) into a one-row-per-garment wide layout so the HF dataset viewer shows every attribute — three images plus 25 metadata columns — on a single row. Previous v3 releases stored one row per (garment, view) with satellite tables that had to be joined manually. That layout is preserved in git history if you need it; the… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/zenodo-second-hand-fashion-v3.imageimage-classification10K<n<100K2 likes530 downloads23d agoHugging Face30yuzhous /lekiwi_second_floor_0916_environmentThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "lekiwi", "total_episodes": 50, "total_frames": 15084, "total_tasks": 1, "total_videos": 50, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yuzhous/lekiwi_second_floor_0916_environment.tabularrobotics10K<n<100K0 likes516 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.