CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AiAF /SCPWiki-Cleaned-PDF-Archivesdocumenttext-generationn<1K1 likes7.7k downloads1y agoHugging Face02common-pile /github_archive GitHub Archive Description According to GitHub’s terms of service, issues and pull request descriptions—along with the their comments—inherit the license of their associated repository. To collect this data, we used the GitHub Archive’s public BigQuery table of events to extracted all issue, pull request, and comment events since 2011 and aggregated them into threads. The table appeared to be missing “edit” events so the text from each comment is the original from when… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive.texttext-generation10M<n<100M2 likes4.3k downloads1y agoHugging Face03common-pile /github_archive_filtered GitHub Archive Description According to GitHub’s terms of service, issues and pull request descriptions—along with their comments—inherit the license of their associated repository. To collect this data, we used the GitHub Archive’s public BigQuery table of events to extract all issue, pull request, and comment events since 2011 and aggregated them into threads. The table appeared to be missing “edit” events so the text from each comment is the original from when it was… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive_filtered.texttext-generation10M<n<100M2 likes695 downloads1y agoHugging Face04aurora-m /biden-harris-redteam-archived THIS IS AN ARCHIVED VERSION Biden-Harris Redteam: A red-teaming dataset focusing on the Biden-Harris AI Executive Order Dataset Description While building Large Language Models (LLMs), it is crucial to protect them against attacks that could bypass safety guardrails and break their guiding principles. Specifically, LLMs should never generate content promoting or normalizing harmful, illegal, or unethical behavior that may contribute to the harm of the… See the full description on the dataset page: https://huggingface.co/datasets/aurora-m/biden-harris-redteam-archived.texttext-generation10K<n<100K7 likes100 downloads1y agoHugging Face05akumalondon /Rail_Freight_Logistics_Company_Email_Archive_Sample Ukrainian Rail-Freight Correspondence Corpus (Sample) Real operational correspondence from a working freight forwarding business, and the documents attached to it — consignment notes, service acts, invoices, wagon manifests. Not scraped, not synthetic, and never published anywhere before. This is a de-identified sample released for evaluation. It is drawn from a larger private archive; see Full archive below. Published by Akuma London · akumalondon.com Why this… See the full description on the dataset page: https://huggingface.co/datasets/akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample.tabulartext-generation1K<n<10K0 likes90 downloads12d agoHugging Face06ChipHolmes /threat-intelligence-dataset-archive Cyber Threat Intelligence Dataset for LLM Fine-Tuning An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on. The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/threat-intelligence-dataset-archive.texttext-generation10K<n<100K1 likes80 downloads2mo agoHugging Face07Banodoco /discord-archive Discord Archive This is an archive of messages from the Banodoco Discord community, where technical and artistic practitioners have been discussing open source AI art for the past three years. The archive captures a long-running community record of people learning, training, evaluating, and using open source AI art models in practice. It contains discussion around model releases, workflows, tooling, troubleshooting, creative experiments, training details, and the many small… See the full description on the dataset page: https://huggingface.co/datasets/Banodoco/discord-archive.tabulartext-generation1M<n<10M4 likes71 downloads4mo agoHugging Face08croqaz /Synthetic-archive The Synthetic Archive Synthetic Archive is a large synthetic English-text dataset generated from OCR-derived historical and period-style passages. Knowledge cutoff is year 1900. Each source passage was divided into chunks and processed through several generation tasks, including: generating continuations of unfinished passages; creating question-and-answer pairs; extracting and reformulating factual knowledge; rewriting material as a narrative; transforming source material into… See the full description on the dataset page: https://huggingface.co/datasets/croqaz/Synthetic-archive.texttext-generation10M<n<100M0 likes63 downloads1mo agoHugging Face09Wilhelm-Foundation /rare-archive-synthetic-patients Rare Archive Synthetic Patients — SFT Training Data 12,984 synthetic rare disease patient vignettes generated from Orphanet disease profiles. Designed for supervised fine-tuning (SFT) of diagnostic AI models. Part of the Rare AI Archive. All patients are computationally generated. Zero real patient data. Zero PHI. This dataset contains no Protected Health Information. Every vignette is synthetically generated from public Orphanet disease profiles using frequency-weighted phenotype… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelm-Foundation/rare-archive-synthetic-patients.texttext-generation10K<n<100K1 likes60 downloads6mo agoHugging Face10Wilhelm-Foundation /rare-archive-eval-rarearena-rds RareArena RDS — Rare Disease Specialists Evaluation Benchmark 8,562 clinical vignettes across 4,000+ rare diseases for evaluating AI diagnostic reasoning. Part of the Rare AI Archive. Research use only. This dataset is an evaluation benchmark for AI systems. It is NOT intended for clinical decision-making and should NOT be used as a diagnostic tool. Ecosystem Context This evaluation benchmark measures how well models handle the diagnostic reasoning patterns that… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelm-Foundation/rare-archive-eval-rarearena-rds.texttext-generation1K<n<10K1 likes54 downloads6mo agoHugging Face11ChipHolmes /All-CVE-Records-Training-Dataset-archive CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/All-CVE-Records-Training-Dataset-archive.texttext-generation100K<n<1M1 likes53 downloads2mo agoHugging Face12ChipHolmes /Trendyol-Cybersecurity-Instruction-Tuning-Dataset-archive Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/Trendyol-Cybersecurity-Instruction-Tuning-Dataset-archive.texttext-generation10K<n<100K1 likes45 downloads2mo agoHugging Face13nyuuzyou /nopaste-paefchen-archive Dataset Card for Nopaste Paefchen Archive Dataset Summary This dataset is an archive of posts from nopaste.paefchen.net, a now-defunct pastebin-like service. It includes approximately 1.7 million unique posts, identified by sequential IDs starting from 1. The content spans various types of text data, including plain text, formatted text, URLs, and potentially code snippets or other formats in multiple languages. Dataset Structure Data Fields This… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/nopaste-paefchen-archive.texttext-generation1M<n<10M0 likes37 downloads2y agoHugging Face14Wilhelm-Foundation /rare-archive-eval-rarearena-rdc RareArena RDC — Rare Disease Cases Evaluation Benchmark 4,376 clinical vignettes with laboratory test results across rare diseases for evaluating AI diagnostic reasoning with lab data. Part of the Rare AI Archive. Research use only. This dataset is an evaluation benchmark for AI systems. It is NOT intended for clinical decision-making and should NOT be used as a diagnostic tool. How RDC Differs from RDS Feature RDS RDC Records 8,562 4,376 Lab results No… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelm-Foundation/rare-archive-eval-rarearena-rdc.texttext-generation1K<n<10K1 likes36 downloads6mo agoHugging Face15ChipHolmes /Cybersecurity-Dataset-Fenrir-v2.1-archive Cybersecurity Defense Instruction-Tuning Dataset (v2.1) Created by Alican Kiraz TL;DR A ready-to-train dataset of 99,870 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training. Apache-2.0 licensed and production-ready. Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security.… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/Cybersecurity-Dataset-Fenrir-v2.1-archive.texttext-generation10K<n<100K0 likes29 downloads2mo agoHugging Face16mlx-community /tnc-archive Paraacademic institution's archived educational activities metadata scraped and packed as a dataset. Dataset Details Dataset Description Scraped titles and summarized descriptions of the "non-members available" data of the Seminars of The New Centre for Research & Practice, took this from our website where i have a status of god of FireStoreStoNe (FSSN). Regarding the latter, should i not scrap the data with an access to the direct descriptions links? 100%… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/tnc-archive.textfeature-extractionn<1K1 likes25 downloads5mo agoHugging Face17mstyslavity /tnc-archive Paraacademic institution's archived educational activities metadata scraped and packed as a dataset. Dataset Details Dataset Description Scraped titles and summarized descriptions of the "non-members available" data of the Seminars of The New Centre for Research & Practice, took this from our website where i have a status of god of FireStoreStoNe (FSSN). Regarding the latter, should i not scrap the data with an access to the direct descriptions links? 100%… See the full description on the dataset page: https://huggingface.co/datasets/mstyslavity/tnc-archive.textfeature-extractionn<1K0 likes24 downloads5mo agoHugging Face18hunterbown /bell-labs-technical-archive Bell Labs Documents and Stuff This is a conservative public-release subset of the internal BELLA continued-pretraining corpus. It keeps the Bell-system technical material that survived a stricter final pass for public dataset hosting and removes records that still looked risky, off-scope, or too low-signal for a Hugging Face corpus listing. What is in the release Split Documents train 1220 validation 29 test 42 The release contains 1291 documents out… See the full description on the dataset page: https://huggingface.co/datasets/hunterbown/bell-labs-technical-archive.tabulartext-generation1K<n<10K0 likes15 downloads6mo agoHugging Face19nazrinburz /internet_archive_azerbaijanitexttext-generationn<1K0 likes8 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.