datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SCPWiki-Cleaned-PDF-Archivesgithub_archive
GitHub Archive
Description
According to GitHub’s terms of service, issues and pull request descriptions—along with the their comments—inherit the license of their associated repository.
To collect this data, we used the GitHub Archive’s public BigQuery table of events to extracted all issue, pull request, and comment events since 2011 and aggregated them into threads.
The table appeared to be missing “edit” events so the text from each comment is the original from when… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive.google-code-archive
Google Code Archive Dataset
Dataset Description
This dataset was compiled from the Google Code Archive, a preserved snapshot of projects hosted on Google Code, Google's open-source project hosting service that operated from 2006 to 2016. Google Code was one of the major code hosting platforms of its era, hosting hundreds of thousands of open-source projects before its shutdown. The archive provides a unique historical record of open-source development during a formative… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/google-code-archive.github_archive_filtered
GitHub Archive
Description
According to GitHub’s terms of service, issues and pull request descriptions—along with their comments—inherit the license of their associated repository.
To collect this data, we used the GitHub Archive’s public BigQuery table of events to extract all issue, pull request, and comment events since 2011 and aggregated them into threads.
The table appeared to be missing “edit” events so the text from each comment is the original from when it was… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive_filtered.google-code-archive
Google Code Archive Dataset
Dataset Description
This dataset was compiled from the Google Code Archive, a preserved snapshot of projects hosted on Google Code, Google's open-source project hosting service that operated from 2006 to 2016. Google Code was one of the major code hosting platforms of its era, hosting hundreds of thousands of open-source projects before its shutdown. The archive provides a unique historical record of open-source development during a formative… See the full description on the dataset page: https://huggingface.co/datasets/Carrillo16/google-code-archive.google-code-archive
Google Code Archive Dataset
Dataset Description
This dataset was compiled from the Google Code Archive, a preserved snapshot of projects hosted on Google Code, Google's open-source project hosting service that operated from 2006 to 2016. Google Code was one of the major code hosting platforms of its era, hosting hundreds of thousands of open-source projects before its shutdown. The archive provides a unique historical record of open-source development during a formative… See the full description on the dataset page: https://huggingface.co/datasets/0xzanuee/google-code-archive.MoeGirlPedia_wikitext_raw_archiveGlad to see models and datasets were inspired from this dataset, thanks to all who are using this dataset in their training materials.
Feel free to re-upload the contents to places like the Internet Archive (Please follow the license and keep these files as-is) to help preserve this digital asset.
Looking forward to see more models and synthetic datasets trained from this raw archive, good luck!
Note: Due to the content censorship system introduced by MGP on 2024/03/29, it is unclear that… See the full description on the dataset page: https://huggingface.co/datasets/milashkaarshif/MoeGirlPedia_wikitext_raw_archive.google-code-archive
Google Code Archive Dataset
Dataset Description
This dataset was compiled from the Google Code Archive, a preserved snapshot of projects hosted on Google Code, Google's open-source project hosting service that operated from 2006 to 2016. Google Code was one of the major code hosting platforms of its era, hosting hundreds of thousands of open-source projects before its shutdown. The archive provides a unique historical record of open-source development during a formative… See the full description on the dataset page: https://huggingface.co/datasets/Mgmgrand420/google-code-archive.RewardLens-phase2-archive
RewardLens Phase II Archive
This is the final clean Hugging Face evidence archive for the completed RewardLens Phase II eight-model experiment.
What this archive contains
8-model experiment evidence
static judgments
audit judgments
Best-of-N pair graphs
selections
final metrics
analysis
figures/tables
manifests
provenance
validity metadata and frozen annotation materials where available
reproducibility metadata and checksums
Models… See the full description on the dataset page: https://huggingface.co/datasets/jlai300/RewardLens-phase2-archive.securecode-web-archive
SecureCode Web: Traditional Web & Application Security Dataset
Production-grade web security vulnerability dataset with complete incident grounding, 4-turn conversational structure, and comprehensive operational guidance
Paper | GitHub | Dataset | Model Collection | Blog Post
What's new in v2.6
v2.6 restores proper Express.js coverage for the topics whose examples were removed in v2.5.1 (they had
shared one reused answer). 29 new, genuinely distinct Express.js… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/securecode-web-archive.ARCHIVE-TEXT-URLS
Internet Archive English Text URLs Dataset
Dataset Description
This dataset contains 11,151,637 direct download URLs to OCR-processed text files from the Internet Archive's digital library. All entries are English-language texts spanning books, documents, historical records, and various other written materials.
Dataset Summary
Total Rows: 11,151,637
Language: English
Source: Internet Archive
Format: CSV with metadata and direct text file URLs
Text… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/ARCHIVE-TEXT-URLS.archiveii
ArchiveII
ArchiveII is a dataset of RNA sequences and their secondary structures, widely used in RNA secondary structure prediction benchmarks.
ArchiveII contains 2975 RNA samples across 10 RNA families, with sequence lengths ranging from 28 to 2968 nucleotides.
This dataset is frequently used to evaluate RNA secondary structure prediction methods, including those that handle both pseudoknotted and non-pseudoknotted structures.
It is considered complementary to the RNAStrAlign… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/archiveii.lg-longtail-data-selection-experiment-archive-20260829
LG Long-tail data-selection experiment archive
This public repository is the canonical, self-contained archive for the ConvFinQA
data-selection budget-scaling experiments started on 2026-08-29 and the preceding
diversity reproduction run started on 2026-08-26. It replaces the earlier split
model/subset repositories.
The archive preserves the local directory trees in full: selected subsets,
selection manifests, LoRA adapters, DeepSpeed optimizer states, trainer states,
logs, raw… See the full description on the dataset page: https://huggingface.co/datasets/Jongbin-kr/lg-longtail-data-selection-experiment-archive-20260829.google-code-archive
Google Code Archive Dataset
Dataset Description
This dataset was compiled from the Google Code Archive, a preserved snapshot of projects hosted on Google Code, Google's open-source project hosting service that operated from 2006 to 2016. Google Code was one of the major code hosting platforms of its era, hosting hundreds of thousands of open-source projects before its shutdown. The archive provides a unique historical record of open-source development during a formative… See the full description on the dataset page: https://huggingface.co/datasets/NarsAI/google-code-archive.taln-archivesTALN Archives benchmark dataset for keyphrase extraction an generation.biden-harris-redteam-archived
THIS IS AN ARCHIVED VERSION
Biden-Harris Redteam: A red-teaming dataset focusing on the Biden-Harris AI Executive Order
Dataset Description
While building Large Language Models (LLMs), it is crucial to protect them against attacks that could bypass safety guardrails and break their guiding principles. Specifically, LLMs should never generate content promoting or normalizing harmful, illegal, or unethical behavior that may contribute to the harm of the… See the full description on the dataset page: https://huggingface.co/datasets/aurora-m/biden-harris-redteam-archived.sayit-archive-tw
Audrey Tang Transcript Corpus
Public transcripts of Audrey Tang (唐鳳) — 2025 Right Livelihood Laureate, civic hacker, and Taiwan's Cyber Ambassador.
Tang is co-author of Plurality: The Future of Collaborative Technology and Democracy and an inaugural Senior Accelerator Fellow at the Oxford Institute for Ethics in AI. She served as Taiwan's first Digital Minister (2016–2024) and the world's first nonbinary cabinet minister, awarded the Right Livelihood Award for "advancing the social… See the full description on the dataset page: https://huggingface.co/datasets/audreyt/sayit-archive-tw.Rail_Freight_Logistics_Company_Email_Archive_Sample
Ukrainian Rail-Freight Correspondence Corpus (Sample)
Real operational correspondence from a working freight forwarding business, and the
documents attached to it — consignment notes, service acts, invoices, wagon
manifests. Not scraped, not synthetic, and never published anywhere before.
This is a de-identified sample released for evaluation. It is drawn from a larger
private archive; see Full archive below.
Published by Akuma London · akumalondon.com
Why this… See the full description on the dataset page: https://huggingface.co/datasets/akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample.threat-intelligence-dataset-archive
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on.
The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/threat-intelligence-dataset-archive.ARCHIVED-cognitive-states
⚠️ ARCHIVED (2026-06-27)
This dataset is archived (renamed from cds-jb/cognitive-states to cds-jb/ARCHIVED-cognitive-states).
Why. A black-box leakage audit (Qwen3.6-35B text monitor; n=1100 balanced over the 11 components × explicit/implicit; 2-way answer-vs-distractor forced choice, order-randomized) found the recognition task is trivially solvable from text alone — 99.5% accuracy (chance = 50%), so an activation oracle has essentially no headroom to demonstrate value as a… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/ARCHIVED-cognitive-states.elektor-archive-dataset
⚡ Elektor Magazine Electronics & Embedded Systems Dataset (1975–2022)
A comprehensive, high-quality instruction-tuning, preference optimization (DPO), and RAG dataset compiled from 11,700+ articles published in Elektor Magazine between 1975 and 2022.
This dataset covers analog/digital circuit design, microcontrollers (AVR, PIC, ESP32, STM32, ARM), RF/communications, power electronics, test & measurement equipment, and audio engineering.
⚠️ Important Disclaimers &… See the full description on the dataset page: https://huggingface.co/datasets/onkanat/elektor-archive-dataset.HoloMol-Pretrain-Archive
HoloMol Pretraining Source Archive
This public dataset repository preserves source snapshots used in the HoloMol
pretraining data pipeline, together with source attribution and recovery
information. It is being populated incrementally.
Swiss-Prot source snapshot
The Swiss-Prot pilot snapshot is listed below. This repository is
not the complete HoloMol pretraining corpus, nor a completed archival backup.
Path
Contents
File size… See the full description on the dataset page: https://huggingface.co/datasets/Csyxx/HoloMol-Pretrain-Archive.worm-chain-archive
WORM Chain Archive
Append-only SHA-256 audit chain logs from sovereign agent execution across the SnapKitty stack. Every entry is cryptographically sealed — no record can be modified or deleted after creation.
Files
File
Source
Description
agentscope-sift.jsonl
agentscope-sift
Security forensic triage execution chain
apl-shell-chain.jsonl
all-apl
APL shell execution WORM records
bob-voyager.jsonl
bob-voyager
BOB agent exploration chain… See the full description on the dataset page: https://huggingface.co/datasets/Snapkitty/worm-chain-archive.discord-archive
Discord Archive
This is an archive of messages from the Banodoco Discord community, where
technical and artistic practitioners have been discussing open source AI art for
the past three years.
The archive captures a long-running community record of people learning,
training, evaluating, and using open source AI art models in practice. It
contains discussion around model releases, workflows, tooling, troubleshooting,
creative experiments, training details, and the many small… See the full description on the dataset page: https://huggingface.co/datasets/Banodoco/discord-archive.sefd-archive-100k-analysis-sample-qwen3-20260524
SEFD Archive 100k Analysis Sample Qwen3 20260524
Retained artifacts for the completed archive-wide 100,000-filing Stanford EDGAR Filings Dataset (SEFD) analysis sample used in the arXiv paper update. The sample contains 2,971,490,909 final SEFD tokens, counted with the Qwen3-1.7B tokenizer.
This repository is a new versioned artifact and intentionally does not replace the earlier sfd-archive-100k-analysis-sample repository used for the original conference submission.
Included:… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/sefd-archive-100k-analysis-sample-qwen3-20260524.ms-codeplex-archive
Microsoft CodePlex Archive Dataset
Dataset Description
Source code from the Microsoft CodePlex Archive on the Internet Archive. CodePlex was Microsoft's open-source project hosting service from 2006 to 2017, popular for .NET and Windows projects.
Dataset Summary
Statistic
Value
Total Files
5,043,730
Total Repositories
38,087
Total Size
3.6 GB (compressed Parquet)
Programming Languages
91
File Format
Parquet with Zstd compression (10 files)… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/ms-codeplex-archive.uyghur-archive
🌟 Uyghur AI Corpus: Bridging Heritage & Technology
ئۇيغۇرچە سۈنئىي ئىدراك خەزىنىسى: مىراس ۋە تېخنىكا كۆۋرۈكى
🌹 Overview / ئومۇمىي ئۇچۇر
EN: The Uyghur AI Corpus is an initiative to preserve and thrive the Uyghur language in the AI era. It serves as a foundational resource for training Large Language Models (LLMs) to understand, generate, and translate Uyghur with high proficiency.
UG: بۇ ئامبار — ئۇيغۇر تىلىنىڭ رەقەملىك دۇنيادىكى ئورنىنى… See the full description on the dataset page: https://huggingface.co/datasets/Uyghur-Corpus/uyghur-archive.Synthetic-archive
The Synthetic Archive
Synthetic Archive is a large synthetic English-text dataset generated from OCR-derived historical and period-style passages. Knowledge cutoff is year 1900.
Each source passage was divided into chunks and processed through several generation tasks, including:
generating continuations of unfinished passages;
creating question-and-answer pairs;
extracting and reformulating factual knowledge;
rewriting material as a narrative;
transforming source material into… See the full description on the dataset page: https://huggingface.co/datasets/croqaz/Synthetic-archive.rare-archive-synthetic-patients
Rare Archive Synthetic Patients — SFT Training Data
12,984 synthetic rare disease patient vignettes generated from Orphanet disease profiles. Designed for supervised fine-tuning (SFT) of diagnostic AI models. Part of the Rare AI Archive.
All patients are computationally generated. Zero real patient data. Zero PHI.
This dataset contains no Protected Health Information. Every vignette is synthetically generated from public Orphanet disease profiles using frequency-weighted phenotype… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelm-Foundation/rare-archive-synthetic-patients.rare-archive-eval-rarearena-rds
RareArena RDS — Rare Disease Specialists Evaluation Benchmark
8,562 clinical vignettes across 4,000+ rare diseases for evaluating AI diagnostic reasoning. Part of the Rare AI Archive.
Research use only. This dataset is an evaluation benchmark for AI systems. It is NOT intended for clinical decision-making and should NOT be used as a diagnostic tool.
Ecosystem Context
This evaluation benchmark measures how well models handle the diagnostic reasoning patterns that… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelm-Foundation/rare-archive-eval-rarearena-rds.
