datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
govdocs1-pdf-source
govdocs1: source PDF files
[!NOTE]
Converted versions of other document types (word, txt, etc) are available in this repo
This is ~220,000 open-access PDF documents (about 6.6M pages) from the dataset govdocs1. It wants to be OCR'd.
Uploaded as tar file pieces of ~10 GiB each due to size/file count limits with an index.csv covering details
5,000 randomly sampled PDFs are available unarchived in sample/. Hugging Face supports previewing these in-browser, for example this one… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-pdf-source.eiken-pdfTechie_Raw_PDFpdf-parsing-bench-resultsICLR-pdfs
Dataset Card for "ICLR-pdfs"
More Information needed
unboxgov-khmdhs-pdfs
KHMDHS Contract Attachment PDFs
Per-contract attachment PDFs from the Greek public-procurement registry
KHMDHS, downsampled with Ghostscript /ebook (~150 DPI) for fast viewing
and compact storage. Files are packed into per-month SQLite shards
(pdfs/<YYYY-MM>.sqlite, one row per ΑΔΑΜ, keyed by referenceNumber).
These are lossy, downsampled copies. The authoritative originals remain at
KHMDHS: https://cerpp.eprocurement.gov.gr/khmdhs-opendata/contract/attachment/{ΑΔΑΜ}.
This… See the full description on the dataset page: https://huggingface.co/datasets/vasilisplavos/unboxgov-khmdhs-pdfs.pdf_science_questions_verified_r1_traces__2_24_25
Dataset card for pdf_science_questions_verified_r1_traces__2_24_25
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"url": "https://www.ttcho.com/_files/ugd/988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf",
"filename": "988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf",
"success": true,
"page_count": 37,
"page_number": 1,
"question_choices_solutions": "QUESTION: What is the identity of X in the reaction 14N + 1n \u2192… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/pdf_science_questions_verified_r1_traces__2_24_25.PDF_and_SCP_unfiltered_organic_chemistry_questionsHARMLESS_Synthetic_Injected_PDFs_EDA
Injected PDFs - EDA and Evaluation Corpus
This repository holds the exploratory data analysis for a project on detecting harmless-but-real
attack payloads injected into PDF files, together with the dataset that analysis produced.
The project has two halves, both in the notebook Final_project_V7_EDA.ipynb:
Question
Input
Part 1
Is our synthetic corpus a stand-in for real malware, or is it something else?
The published CIC feature table (11,126 x 34)
Part 2
Is our… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/HARMLESS_Synthetic_Injected_PDFs_EDA.Pashto-Textbooks-PDFs-Corpus
Pashto Textbooks and PDFs Corpus
Languages: psLicense: cc-by-4.0Task categories: text-generation, feature-extractionSize categories: n<1K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, feature-extraction tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/Pashto-Textbooks-PDFs-Corpus")
print(dataset)
Configs
default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/Pashto-Textbooks-PDFs-Corpus.pdf_science_questions_verifiable_r1_traces__2_24_25
Dataset card for pdf_science_questions_verifiable_r1_traces__2_24_25
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"url": "https://www.ttcho.com/_files/ugd/988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf",
"filename": "988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf",
"success": true,
"page_count": 37,
"page_number": 1,
"question_choices_solutions": "QUESTION: What is the identity of X in the reaction 14N + 1n \u2192… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/pdf_science_questions_verifiable_r1_traces__2_24_25.ICLR-pdfs-linebreaks
Dataset Card for "ICLR-pdfs-linebreaks"
More Information needed
pdf_qa_r1_annotated_verifiedStenCore-PDF
StenCore — FinePDFs-Edu Curated
By StentorLabs
StenCore is StentorLabs' first dataset release — a quality-first resource built from the ground up for training language models. Derived from HuggingFaceFW/finepdfs-edu, every document passed a 14-stage automated curation pipeline including heuristic filtering, language ID, PII redaction, toxicity screening, eval decontamination, multi-strategy deduplication, KenLM + neural perplexity scoring, and domain reweighting.
⚠️ Privacy &… See the full description on the dataset page: https://huggingface.co/datasets/StentorLabs/StenCore-PDF.Evaluation_of_OpenSource_Models_for_PDF_Injection_Recognition
Injected PDFs - Model Evaluation
This repository holds the model evaluation stage of a project on detecting harmless-but-real
attack payloads injected into PDF files, together with the artefacts it produced for the
application.
Nothing is trained here. Seven off-the-shelf models are measured against the same 1,100 PDFs,
and the two winners are exported for the app to load.
Question
Candidates
Winner
Part A
Which files look like this one?
3 embedding models x 2 inputs… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/Evaluation_of_OpenSource_Models_for_PDF_Injection_Recognition.organic_chemistry_pdf_word_searchimslp-pdf-index
IMSLP PDF Index
This dataset is the canonical PDF-level index for the ReScore IMSLP PDF
collection. It contains one row per unique IMSLP PDF and points to the PDF
payload stored in cminst/imslp-raw-pdf-collection.
The PDF payload repository is append-only and may contain duplicate rows from
retry launches. This index is deduplicated by imslp_id; duplicate content was
validated to have identical SHA256, byte size, and page count before publishing.
Summary
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/cminst/imslp-pdf-index.Coptic-PDF-Corpuspdfsys-page-v2-demo
pdfsys.page/v2 — 格式演示数据集
pdfsys.page/v2 是 pdfsystem_mnbvc
的 L2 发布格式,为 MNBVC 中文语料的 PB 级 PDF 流水线设计。
这是一个格式演示,不是训练语料。 25 页、18 份文档,只够说明 schema 长什么样、
三种视图怎么取。真实语料是 21.8 万份 PDF 的量级。
来源提示:这里的 PDF 页来自 OmniDocBench
与 olmOCR-bench 两个公开
benchmark,逐份的上游许可未经核实。放出来是为了说明数据格式,不是为了再分发这些
文档本身——要拿去用请自行确认源文档的许可。详见文末「来源与许可」。
一句话设计
一行一页,主键 (doc_id, page_index) ——这个身份来自 PDF 本身,不是模型造出来的;
页文本里内联图标记来承载图文交错;模型派生的结构是旁边一列可丢弃的增强;
图像像素要么是裁剪图、要么是整页光栅,二选一。
里面有什么
config
行数… See the full description on the dataset page: https://huggingface.co/datasets/miracleyin/pdfsys-page-v2-demo.pdf-steganalysis-corpus
PDF Steganalysis Corpus v2
A research benchmark for clean-versus-stego detection and method attribution,
plus a separate downloadable demonstration made from original synthetic PDFs.
"Clean" means the original document before our embedding, not malware-free.
This is not a malware collection.
What to download
Files
Purpose
historical-{train,validation,test}.parquet
13,637 verified historical stego records + 1,939 clean originals, with labels and 38… See the full description on the dataset page: https://huggingface.co/datasets/manj0220/pdf-steganalysis-corpus.napierone-pdf-raw
BEE-spoke-data/napierone-pdf-raw
NapierOne PDF files converted with marker.
detected languages
Counter({'en': 4665,
'nl': 2,
'fi': 7,
'fr': 8,
'cy': 54,
'sq': 1,
'it': 1,
'unknown-error': 5,
'sk': 1,
'es': 2,
'de': 3,
'ro': 1,
'pl': 1,
'zh': 1,
'so': 1,
'ml': 1})
organic_chemistry_pdfUN_PDF_RECORD_SET
Dataset Card for "UN_PDF_RECORD_SET"
More Information needed
pdf-upload-caps-2026
PDF upload caps on public portals (2026)
How large can a PDF be before a government, university or job portal rejects it? This dataset records the published file size limit of 162 portals in France, the United States, the United Kingdom, Germany, Spain and India, each with the exact wording of the limit and a link to the official page where it was found.
It was collected in September 2026 for the EasyPDF study The 1 MB Problem: PDF File Size Statistics for 2026 (French version:… See the full description on the dataset page: https://huggingface.co/datasets/EasyPDF/pdf-upload-caps-2026.DeepEval-Question-Answer-Dataset-for-RAG-Evaluation-A2A-And-ACP-PDFpdf-BZU-catalog-for-GPU-projectall_filtered_unverified_pdfs_pipelineimslp-raw-pdf-collection
IMSLP Raw PDF Collection
This dataset stores raw public-domain IMSLP PDF downloads collected by ReScore.
The raw PDF archive is append-only. Each Modal download launch publishes tar
shards under pdf_shards/ and matching JSONL manifests under manifests/.
Default collection id: imslp_pdf_collection.
The coordinator remains the source of claim/download state. Manifest rows link
each tar member back to the IMSLP id, source metadata, PDF checksum, page count,
Modal profile, launch id… See the full description on the dataset page: https://huggingface.co/datasets/cminst/imslp-raw-pdf-collection.pdfa-eng-wds-eval-filteredfiltered_pdf_flow
