datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jfk-archives
Dataset Card for JFK Archives
This dataset is a collection of all records pertaining to the assassination of the
US president, John F. Kennedy, released until April 2025 through archives.org
by the US government.
Dataset Details
Dataset Description
The original data downloaded from archives.org
consists of 56,300 scanned documents in PDF format, released until April 2025. The files are
organized by their release year(s): 2107-2018, 2021, 2022, 2023 and 2025.… See the full description on the dataset page: https://huggingface.co/datasets/farhanhubble/jfk-archives.securecode-web-archive
SecureCode Web: Traditional Web & Application Security Dataset
Production-grade web security vulnerability dataset with complete incident grounding, 4-turn conversational structure, and comprehensive operational guidance
Paper | GitHub | Dataset | Model Collection | Blog Post
What's new in v2.6
v2.6 restores proper Express.js coverage for the topics whose examples were removed in v2.5.1 (they had
shared one reused answer). 29 new, genuinely distinct Express.js… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/securecode-web-archive.ARCHIVE-TEXT-URLS
Internet Archive English Text URLs Dataset
Dataset Description
This dataset contains 11,151,637 direct download URLs to OCR-processed text files from the Internet Archive's digital library. All entries are English-language texts spanning books, documents, historical records, and various other written materials.
Dataset Summary
Total Rows: 11,151,637
Language: English
Source: Internet Archive
Format: CSV with metadata and direct text file URLs
Text… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/ARCHIVE-TEXT-URLS.hotpot_qa_archiveHotpotQA is a new dataset with 113k Wikipedia-based question-answer pairs with four key features:
(1) the questions require finding and reasoning over multiple supporting documents to answer;
(2) the questions are diverse and not constrained to any pre-existing knowledge bases or knowledge schemas;
(3) we provide sentence-level supporting facts required for reasoning, allowingQA systems to reason with strong supervisionand explain the predictions;
(4) we offer a new type of factoid comparison questions to testQA systems’ ability to extract relevant facts and perform necessary comparison.cbi-archive-corpus
Central Bank of Ireland Public Archive Corpus
A page-anchored, provenance-classified corpus of the Central Bank of Ireland's
public document archive. 5,568 documents and 89,242 page or pseudo-page rows.
PDF rows have true source-page anchors; most Office and archive rows do not.
This is an unofficial derived work. It is not published by, affiliated with, or
endorsed by the Central Bank of Ireland.
What makes this different from a pile of scraped PDFs
Two things.… See the full description on the dataset page: https://huggingface.co/datasets/aditya487/cbi-archive-corpus.policystrategies-archive
🏛️ Open-Source Macro-Strategy, Financial History & Intelligence Archive
🌐 Overview & Institutional Mission
This public repository serves as the official open-source knowledge graph and metadata registry for r/policystrategies.
We aggregate, document, and cross-reference declassified historical intelligence dossiers, sovereign debt crises, systemic market manipulations, and geoeconomic conflicts using verified open-source intelligence (OSINT) and primary… See the full description on the dataset page: https://huggingface.co/datasets/stratigahq/policystrategies-archive.sayit-archive-tw
Audrey Tang Transcript Corpus
Public transcripts of Audrey Tang (唐鳳) — 2025 Right Livelihood Laureate, civic hacker, and Taiwan's Cyber Ambassador.
Tang is co-author of Plurality: The Future of Collaborative Technology and Democracy and an inaugural Senior Accelerator Fellow at the Oxford Institute for Ethics in AI. She served as Taiwan's first Digital Minister (2016–2024) and the world's first nonbinary cabinet minister, awarded the Right Livelihood Award for "advancing the social… See the full description on the dataset page: https://huggingface.co/datasets/audreyt/sayit-archive-tw.Rail_Freight_Logistics_Company_Email_Archive_Sample
Ukrainian Rail-Freight Correspondence Corpus (Sample)
Real operational correspondence from a working freight forwarding business, and the
documents attached to it — consignment notes, service acts, invoices, wagon
manifests. Not scraped, not synthetic, and never published anywhere before.
This is a de-identified sample released for evaluation. It is drawn from a larger
private archive; see Full archive below.
Published by Akuma London · akumalondon.com
Why this… See the full description on the dataset page: https://huggingface.co/datasets/akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample.threat-intelligence-dataset-archive
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on.
The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/threat-intelligence-dataset-archive.elektor-archive-dataset
⚡ Elektor Magazine Electronics & Embedded Systems Dataset (1975–2022)
A comprehensive, high-quality instruction-tuning, preference optimization (DPO), and RAG dataset compiled from 11,700+ articles published in Elektor Magazine between 1975 and 2022.
This dataset covers analog/digital circuit design, microcontrollers (AVR, PIC, ESP32, STM32, ARM), RF/communications, power electronics, test & measurement equipment, and audio engineering.
⚠️ Important Disclaimers &… See the full description on the dataset page: https://huggingface.co/datasets/onkanat/elektor-archive-dataset.sefd-archive-100k-analysis-sample-qwen3-20260524
SEFD Archive 100k Analysis Sample Qwen3 20260524
Retained artifacts for the completed archive-wide 100,000-filing Stanford EDGAR Filings Dataset (SEFD) analysis sample used in the arXiv paper update. The sample contains 2,971,490,909 final SEFD tokens, counted with the Qwen3-1.7B tokenizer.
This repository is a new versioned artifact and intentionally does not replace the earlier sfd-archive-100k-analysis-sample repository used for the original conference submission.
Included:… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/sefd-archive-100k-analysis-sample-qwen3-20260524.scriptinghelpers-archiveThis is a small dataset containing question/answer articles on (now defunct) scriptinghelpers.org, which was a StackExchange-like website for Roblox developers. It was scraped from archive.org about a year ago and parsed with regex.
Trivia
There are 2097 archived articles with accepted answers out of 5129.
The earliest question archived was posted on Sat Feb 1 23:57:31 2014 with question_id 1 by TheGuyWithAShortName entitled HttpService - PostAsync . It had an accepted answer… See the full description on the dataset page: https://huggingface.co/datasets/densenet/scriptinghelpers-archive.rare-archive-synthetic-patients
Rare Archive Synthetic Patients — SFT Training Data
12,984 synthetic rare disease patient vignettes generated from Orphanet disease profiles. Designed for supervised fine-tuning (SFT) of diagnostic AI models. Part of the Rare AI Archive.
All patients are computationally generated. Zero real patient data. Zero PHI.
This dataset contains no Protected Health Information. Every vignette is synthetically generated from public Orphanet disease profiles using frequency-weighted phenotype… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelm-Foundation/rare-archive-synthetic-patients.project-23-argentina-travel-archive
🇦🇷 Project 23 Argentina Travel Archive
Dataset Description
This dataset contains a structured archive of Argentina-focused travel, media reference, article, video transcript, and photography metadata records from the Samuel & Audrey Media Network.
The archive is part of Project 23, a long-term effort to document Argentina’s 23 provinces through travel guides, videos, photography, regional logistics, cultural coverage, and public source records. The dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/project-23-argentina-travel-archive.rare-archive-eval-rarearena-rds
RareArena RDS — Rare Disease Specialists Evaluation Benchmark
8,562 clinical vignettes across 4,000+ rare diseases for evaluating AI diagnostic reasoning. Part of the Rare AI Archive.
Research use only. This dataset is an evaluation benchmark for AI systems. It is NOT intended for clinical decision-making and should NOT be used as a diagnostic tool.
Ecosystem Context
This evaluation benchmark measures how well models handle the diagnostic reasoning patterns that… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelm-Foundation/rare-archive-eval-rarearena-rds.top-100-travel-blogs-2010s-archive
Top 100 Travel Blogs 2010s Historical Archive
Historical ranking archive — not a current ranking.
This dataset preserves the Nomadic Samuel Top 100 Travel Blogs ranking from the early-to-mid 2010s as a structured historical archive. It includes the final Top 100 composite ranking, additional composite ranking rows from the source page, metric-specific ranking tables, blog entity records, methodology context, origin-story context, academic/research references, and public references… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/top-100-travel-blogs-2010s-archive.Trendyol-Cybersecurity-Instruction-Tuning-Dataset-archive
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/Trendyol-Cybersecurity-Instruction-Tuning-Dataset-archive.sfd-archive-1b-source-format-sample
SFD Archive 1B-Token Source-Format Sample
Cleaned artifacts for an archive-wide SFD source-format analysis sample. The sanitized filing_stats.jsonl.gz contains 37,534 parsed filing rows and 997,469,365 final SFD tokens. The sampled manifest contains 100,000 candidate rows. summary.json is recomputed from the uploaded filing stats; source_summary_checkpoint.json preserves the original run checkpoint summary. Parser stdout tails, local paths, temporary raw SEC downloads, and process… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/sfd-archive-1b-source-format-sample.securecode-archive
SecureCode: Comprehensive Security Training Dataset for AI Coding Assistants
The largest open security training dataset for AI coding assistants, covering both traditional web security and AI/ML security
Overview
SecureCode combines 2,372 security-focused training examples into a single, unified dataset with HuggingFace configs for flexible loading. Every example provides vulnerable code, explains why it's dangerous, demonstrates a secure alternative, and… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/securecode-archive.rare-archive-eval-rarearena-rdc
RareArena RDC — Rare Disease Cases Evaluation Benchmark
4,376 clinical vignettes with laboratory test results across rare diseases for evaluating AI diagnostic reasoning with lab data. Part of the Rare AI Archive.
Research use only. This dataset is an evaluation benchmark for AI systems. It is NOT intended for clinical decision-making and should NOT be used as a diagnostic tool.
How RDC Differs from RDS
Feature
RDS
RDC
Records
8,562
4,376
Lab results
No… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelm-Foundation/rare-archive-eval-rarearena-rdc.sfd-archive-100k-analysis-sample
SFD Archive 100k Analysis Sample
Retained artifacts for the completed archive-wide 100k-filing SFD analysis sample. The run processed 99,895 filings, produced 99,607 successful parses, and contains 2,698,335,638 final SFD tokens. This repository contains aggregate run metadata, processed accessions, and paper-analysis CSV/JSON metrics. Per-filing markdown outputs and temporary raw SEC downloads are not included.
early-travel-blogging-directory-archive
Early Travel Blogging Directory Archive
Historical directory archive — not a current recommendation list.
DOI: 10.57967/hf/8961
This dataset preserves the Nomadic Samuel Travel Blog Directory / Links page as a structured historical archive of early travel blogs, independent travel websites, backpacking sites, food-and-travel blogs, photography blogs, family travel blogs, digital nomad projects, regional travel guides, and related creator-era web properties listed on the source page.… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/early-travel-blogging-directory-archive.23f-declassified-archive-1981
23-F Papel — Declassified Archive of the Spanish Coup d'État (1981)
165 declassified Spanish government documents about the failed coup d'état of 23 February 1981,
fully structured and bilingual (Spanish / English).
Canonical source: https://23fpapel.es
About the event
The 23-F was a failed coup d'état in Spain on 23 February 1981. Lieutenant Colonel Antonio Tejero
Molina stormed the Congress of Deputies with 200 Civil Guard officers, holding the entire parliament… See the full description on the dataset page: https://huggingface.co/datasets/823409idfdu234/23f-declassified-archive-1981.
