CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AiAF /SCPWiki-Cleaned-PDF-Archivesdocumenttext-generationn<1K1 likes7.7k downloads1y agoHugging Face02common-pile /github_archive GitHub Archive Description According to GitHub’s terms of service, issues and pull request descriptions—along with the their comments—inherit the license of their associated repository. To collect this data, we used the GitHub Archive’s public BigQuery table of events to extracted all issue, pull request, and comment events since 2011 and aggregated them into threads. The table appeared to be missing “edit” events so the text from each comment is the original from when… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive.texttext-generation10M<n<100M2 likes4.3k downloads1y agoHugging Face03nyuuzyou /google-code-archive Google Code Archive Dataset Dataset Description This dataset was compiled from the Google Code Archive, a preserved snapshot of projects hosted on Google Code, Google's open-source project hosting service that operated from 2006 to 2016. Google Code was one of the major code hosting platforms of its era, hosting hundreds of thousands of open-source projects before its shutdown. The archive provides a unique historical record of open-source development during a formative… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/google-code-archive.texttext-generation10M<n<100M73 likes1.7k downloads8mo agoHugging Face04common-pile /github_archive_filtered GitHub Archive Description According to GitHub’s terms of service, issues and pull request descriptions—along with their comments—inherit the license of their associated repository. To collect this data, we used the GitHub Archive’s public BigQuery table of events to extract all issue, pull request, and comment events since 2011 and aggregated them into threads. The table appeared to be missing “edit” events so the text from each comment is the original from when it was… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive_filtered.texttext-generation10M<n<100M2 likes695 downloads1y agoHugging Face05Carrillo16 /google-code-archive Google Code Archive Dataset Dataset Description This dataset was compiled from the Google Code Archive, a preserved snapshot of projects hosted on Google Code, Google's open-source project hosting service that operated from 2006 to 2016. Google Code was one of the major code hosting platforms of its era, hosting hundreds of thousands of open-source projects before its shutdown. The archive provides a unique historical record of open-source development during a formative… See the full description on the dataset page: https://huggingface.co/datasets/Carrillo16/google-code-archive.texttext-generation10M<n<100M0 likes548 downloads8mo agoHugging Face060xzanuee /google-code-archive Google Code Archive Dataset Dataset Description This dataset was compiled from the Google Code Archive, a preserved snapshot of projects hosted on Google Code, Google's open-source project hosting service that operated from 2006 to 2016. Google Code was one of the major code hosting platforms of its era, hosting hundreds of thousands of open-source projects before its shutdown. The archive provides a unique historical record of open-source development during a formative… See the full description on the dataset page: https://huggingface.co/datasets/0xzanuee/google-code-archive.texttext-generation10M<n<100M0 likes503 downloads8mo agoHugging Face07milashkaarshif /MoeGirlPedia_wikitext_raw_archiveGlad to see models and datasets were inspired from this dataset, thanks to all who are using this dataset in their training materials. Feel free to re-upload the contents to places like the Internet Archive (Please follow the license and keep these files as-is) to help preserve this digital asset. Looking forward to see more models and synthetic datasets trained from this raw archive, good luck! Note: Due to the content censorship system introduced by MGP on 2024/03/29, it is unclear that… See the full description on the dataset page: https://huggingface.co/datasets/milashkaarshif/MoeGirlPedia_wikitext_raw_archive.texttext-generation100K<n<1M38 likes490 downloads8mo agoHugging Face08Mgmgrand420 /google-code-archive Google Code Archive Dataset Dataset Description This dataset was compiled from the Google Code Archive, a preserved snapshot of projects hosted on Google Code, Google's open-source project hosting service that operated from 2006 to 2016. Google Code was one of the major code hosting platforms of its era, hosting hundreds of thousands of open-source projects before its shutdown. The archive provides a unique historical record of open-source development during a formative… See the full description on the dataset page: https://huggingface.co/datasets/Mgmgrand420/google-code-archive.texttext-generation10M<n<100M0 likes459 downloads8mo agoHugging Face09jlai300 /RewardLens-phase2-archive RewardLens Phase II Archive This is the final clean Hugging Face evidence archive for the completed RewardLens Phase II eight-model experiment. What this archive contains 8-model experiment evidence static judgments audit judgments Best-of-N pair graphs selections final metrics analysis figures/tables manifests provenance validity metadata and frozen annotation materials where available reproducibility metadata and checksums Models… See the full description on the dataset page: https://huggingface.co/datasets/jlai300/RewardLens-phase2-archive.visual-question-answering0 likes374 downloads9d agoHugging Face10ChipHolmes /securecode-web-archive SecureCode Web: Traditional Web & Application Security Dataset Production-grade web security vulnerability dataset with complete incident grounding, 4-turn conversational structure, and comprehensive operational guidance Paper | GitHub | Dataset | Model Collection | Blog Post What's new in v2.6 v2.6 restores proper Express.js coverage for the topics whose examples were removed in v2.5.1 (they had shared one reused answer). 29 new, genuinely distinct Express.js… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/securecode-web-archive.texttext-generation1K<n<10K0 likes329 downloads2mo agoHugging Face11Navanjana /ARCHIVE-TEXT-URLS Internet Archive English Text URLs Dataset Dataset Description This dataset contains 11,151,637 direct download URLs to OCR-processed text files from the Internet Archive's digital library. All entries are English-language texts spanning books, documents, historical records, and various other written materials. Dataset Summary Total Rows: 11,151,637 Language: English Source: Internet Archive Format: CSV with metadata and direct text file URLs Text… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/ARCHIVE-TEXT-URLS.texttext-generation1M<n<10M1 likes297 downloads10mo agoHugging Face12multimolecule /archiveii ArchiveII ArchiveII is a dataset of RNA sequences and their secondary structures, widely used in RNA secondary structure prediction benchmarks. ArchiveII contains 2975 RNA samples across 10 RNA families, with sequence lengths ranging from 28 to 2968 nucleotides. This dataset is frequently used to evaluate RNA secondary structure prediction methods, including those that handle both pseudoknotted and non-pseudoknotted structures. It is considered complementary to the RNAStrAlign… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/archiveii.texttext-generation1K<n<10K0 likes284 downloads2mo agoHugging Face13Jongbin-kr /lg-longtail-data-selection-experiment-archive-20260829 LG Long-tail data-selection experiment archive This public repository is the canonical, self-contained archive for the ConvFinQA data-selection budget-scaling experiments started on 2026-08-29 and the preceding diversity reproduction run started on 2026-08-26. It replaces the earlier split model/subset repositories. The archive preserves the local directory trees in full: selected subsets, selection manifests, LoRA adapters, DeepSpeed optimizer states, trainer states, logs, raw… See the full description on the dataset page: https://huggingface.co/datasets/Jongbin-kr/lg-longtail-data-selection-experiment-archive-20260829.text-generation0 likes240 downloads1d agoHugging Face14NarsAI /google-code-archive Google Code Archive Dataset Dataset Description This dataset was compiled from the Google Code Archive, a preserved snapshot of projects hosted on Google Code, Google's open-source project hosting service that operated from 2006 to 2016. Google Code was one of the major code hosting platforms of its era, hosting hundreds of thousands of open-source projects before its shutdown. The archive provides a unique historical record of open-source development during a formative… See the full description on the dataset page: https://huggingface.co/datasets/NarsAI/google-code-archive.texttext-generation10M<n<100M1 likes137 downloads8mo agoHugging Face15taln-ls2n /taln-archivesTALN Archives benchmark dataset for keyphrase extraction an generation.texttext-generation1K<n<10K3 likes103 downloads4y agoHugging Face16aurora-m /biden-harris-redteam-archived THIS IS AN ARCHIVED VERSION Biden-Harris Redteam: A red-teaming dataset focusing on the Biden-Harris AI Executive Order Dataset Description While building Large Language Models (LLMs), it is crucial to protect them against attacks that could bypass safety guardrails and break their guiding principles. Specifically, LLMs should never generate content promoting or normalizing harmful, illegal, or unethical behavior that may contribute to the harm of the… See the full description on the dataset page: https://huggingface.co/datasets/aurora-m/biden-harris-redteam-archived.texttext-generation10K<n<100K7 likes100 downloads1y agoHugging Face17audreyt /sayit-archive-tw Audrey Tang Transcript Corpus Public transcripts of Audrey Tang (唐鳳) — 2025 Right Livelihood Laureate, civic hacker, and Taiwan's Cyber Ambassador. Tang is co-author of Plurality: The Future of Collaborative Technology and Democracy and an inaugural Senior Accelerator Fellow at the Oxford Institute for Ethics in AI. She served as Taiwan's first Digital Minister (2016–2024) and the world's first nonbinary cabinet minister, awarded the Right Livelihood Award for "advancing the social… See the full description on the dataset page: https://huggingface.co/datasets/audreyt/sayit-archive-tw.tabulartext-generation100K<n<1M0 likes90 downloads7mo agoHugging Face18akumalondon /Rail_Freight_Logistics_Company_Email_Archive_Sample Ukrainian Rail-Freight Correspondence Corpus (Sample) Real operational correspondence from a working freight forwarding business, and the documents attached to it — consignment notes, service acts, invoices, wagon manifests. Not scraped, not synthetic, and never published anywhere before. This is a de-identified sample released for evaluation. It is drawn from a larger private archive; see Full archive below. Published by Akuma London · akumalondon.com Why this… See the full description on the dataset page: https://huggingface.co/datasets/akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample.tabulartext-generation1K<n<10K0 likes90 downloads12d agoHugging Face19ChipHolmes /threat-intelligence-dataset-archive Cyber Threat Intelligence Dataset for LLM Fine-Tuning An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on. The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/threat-intelligence-dataset-archive.texttext-generation10K<n<100K1 likes80 downloads2mo agoHugging Face20cds-jb /ARCHIVED-cognitive-states ⚠️ ARCHIVED (2026-06-27) This dataset is archived (renamed from cds-jb/cognitive-states to cds-jb/ARCHIVED-cognitive-states). Why. A black-box leakage audit (Qwen3.6-35B text monitor; n=1100 balanced over the 11 components × explicit/implicit; 2-way answer-vs-distractor forced choice, order-randomized) found the recognition task is trivially solvable from text alone — 99.5% accuracy (chance = 50%), so an activation oracle has essentially no headroom to demonstrate value as a… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/ARCHIVED-cognitive-states.texttext-generation1M<n<10M0 likes79 downloads3mo agoHugging Face21onkanat /elektor-archive-dataset ⚡ Elektor Magazine Electronics & Embedded Systems Dataset (1975–2022) A comprehensive, high-quality instruction-tuning, preference optimization (DPO), and RAG dataset compiled from 11,700+ articles published in Elektor Magazine between 1975 and 2022. This dataset covers analog/digital circuit design, microcontrollers (AVR, PIC, ESP32, STM32, ARM), RF/communications, power electronics, test & measurement equipment, and audio engineering. ⚠️ Important Disclaimers &… See the full description on the dataset page: https://huggingface.co/datasets/onkanat/elektor-archive-dataset.texttext-generation100K<n<1M0 likes77 downloads2mo agoHugging Face22Csyxx /HoloMol-Pretrain-Archive HoloMol Pretraining Source Archive This public dataset repository preserves source snapshots used in the HoloMol pretraining data pipeline, together with source attribution and recovery information. It is being populated incrementally. Swiss-Prot source snapshot The Swiss-Prot pilot snapshot is listed below. This repository is not the complete HoloMol pretraining corpus, nor a completed archival backup. Path Contents File size… See the full description on the dataset page: https://huggingface.co/datasets/Csyxx/HoloMol-Pretrain-Archive.text-generation0 likes75 downloads6d agoHugging Face23Snapkitty /worm-chain-archive WORM Chain Archive Append-only SHA-256 audit chain logs from sovereign agent execution across the SnapKitty stack. Every entry is cryptographically sealed — no record can be modified or deleted after creation. Files File Source Description agentscope-sift.jsonl agentscope-sift Security forensic triage execution chain apl-shell-chain.jsonl all-apl APL shell execution WORM records bob-voyager.jsonl bob-voyager BOB agent exploration chain… See the full description on the dataset page: https://huggingface.co/datasets/Snapkitty/worm-chain-archive.text-generationn<1K0 likes73 downloads21d agoHugging Face24Banodoco /discord-archive Discord Archive This is an archive of messages from the Banodoco Discord community, where technical and artistic practitioners have been discussing open source AI art for the past three years. The archive captures a long-running community record of people learning, training, evaluating, and using open source AI art models in practice. It contains discussion around model releases, workflows, tooling, troubleshooting, creative experiments, training details, and the many small… See the full description on the dataset page: https://huggingface.co/datasets/Banodoco/discord-archive.tabulartext-generation1M<n<10M4 likes71 downloads4mo agoHugging Face25sfd-anonymous /sefd-archive-100k-analysis-sample-qwen3-20260524 SEFD Archive 100k Analysis Sample Qwen3 20260524 Retained artifacts for the completed archive-wide 100,000-filing Stanford EDGAR Filings Dataset (SEFD) analysis sample used in the arXiv paper update. The sample contains 2,971,490,909 final SEFD tokens, counted with the Qwen3-1.7B tokenizer. This repository is a new versioned artifact and intentionally does not replace the earlier sfd-archive-100k-analysis-sample repository used for the original conference submission. Included:… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/sefd-archive-100k-analysis-sample-qwen3-20260524.text-generation0 likes70 downloads4mo agoHugging Face26nyuuzyou /ms-codeplex-archive Microsoft CodePlex Archive Dataset Dataset Description Source code from the Microsoft CodePlex Archive on the Internet Archive. CodePlex was Microsoft's open-source project hosting service from 2006 to 2017, popular for .NET and Windows projects. Dataset Summary Statistic Value Total Files 5,043,730 Total Repositories 38,087 Total Size 3.6 GB (compressed Parquet) Programming Languages 91 File Format Parquet with Zstd compression (10 files)… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/ms-codeplex-archive.texttext-generation1M<n<10M4 likes65 downloads8mo agoHugging Face27Uyghur-Corpus /uyghur-archive 🌟 Uyghur AI Corpus: Bridging Heritage & Technology ئۇيغۇرچە سۈنئىي ئىدراك خەزىنىسى: مىراس ۋە تېخنىكا كۆۋرۈكى 🌹 Overview / ئومۇمىي ئۇچۇر EN: The Uyghur AI Corpus is an initiative to preserve and thrive the Uyghur language in the AI era. It serves as a foundational resource for training Large Language Models (LLMs) to understand, generate, and translate Uyghur with high proficiency. UG: بۇ ئامبار — ئۇيغۇر تىلىنىڭ رەقەملىك دۇنيادىكى ئورنىنى… See the full description on the dataset page: https://huggingface.co/datasets/Uyghur-Corpus/uyghur-archive.text-generation0 likes64 downloads14d agoHugging Face28croqaz /Synthetic-archive The Synthetic Archive Synthetic Archive is a large synthetic English-text dataset generated from OCR-derived historical and period-style passages. Knowledge cutoff is year 1900. Each source passage was divided into chunks and processed through several generation tasks, including: generating continuations of unfinished passages; creating question-and-answer pairs; extracting and reformulating factual knowledge; rewriting material as a narrative; transforming source material into… See the full description on the dataset page: https://huggingface.co/datasets/croqaz/Synthetic-archive.texttext-generation10M<n<100M0 likes63 downloads1mo agoHugging Face29Wilhelm-Foundation /rare-archive-synthetic-patients Rare Archive Synthetic Patients — SFT Training Data 12,984 synthetic rare disease patient vignettes generated from Orphanet disease profiles. Designed for supervised fine-tuning (SFT) of diagnostic AI models. Part of the Rare AI Archive. All patients are computationally generated. Zero real patient data. Zero PHI. This dataset contains no Protected Health Information. Every vignette is synthetically generated from public Orphanet disease profiles using frequency-weighted phenotype… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelm-Foundation/rare-archive-synthetic-patients.texttext-generation10K<n<100K1 likes60 downloads6mo agoHugging Face30Wilhelm-Foundation /rare-archive-eval-rarearena-rds RareArena RDS — Rare Disease Specialists Evaluation Benchmark 8,562 clinical vignettes across 4,000+ rare diseases for evaluating AI diagnostic reasoning. Part of the Rare AI Archive. Research use only. This dataset is an evaluation benchmark for AI systems. It is NOT intended for clinical decision-making and should NOT be used as a diagnostic tool. Ecosystem Context This evaluation benchmark measures how well models handle the diagnostic reasoning patterns that… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelm-Foundation/rare-archive-eval-rarearena-rds.texttext-generation1K<n<10K1 likes54 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.