datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arxiv-cs-2020-2025-pdfsdom-pi-pdfs-2025
DOM-PI 2025 — PDFs-fonte (Diário Oficial dos Municípios do Piauí)
PDFs originais das publicações de 2025 do Diário Oficial dos Municípios do Piauí,
organizados por Território de Desenvolvimento. São a fonte da qual o corpus textual foi
extraído por OCR/parsing. 41.617 PDFs · ~70 GB.
Dataset de texto derivado (carregável, com limpeza e tiers de qualidade):
gutoportelaa/dom-pi-corpus-2025.
Cobertura de PDFs (parcial): presentes 7 territórios — tabuleiros_alto_parnaiba… See the full description on the dataset page: https://huggingface.co/datasets/gutoportelaa/dom-pi-pdfs-2025.GDP.pdf
GDP.pdf
GDP.pdf measures professional multimodal reasoning over the documents the economy actually runs on: dense reports, contracts, filings, and records in the messy real-world formats professionals work from. It was cited in Anthropic's Fable 5 and Mythos 5 model card.
What it tests
The benchmark contains 100 real-world prompts and PDFs pulled directly from professional workflows across ten domains: Finance, Healthcare, Legal, STEM/Research, Engineering… See the full description on the dataset page: https://huggingface.co/datasets/surgeai/GDP.pdf.20000-meishi-pdfcc-pdf-linkspdfa-eng-wds
Dataset Card for PDF Association dataset (PDFA)
Dataset Summary
PDFA dataset is a document dataset filtered from the SafeDocs corpus, aka CC-MAIN-2021-31-PDF-UNTRUNCATED. The original purpose of that corpus is for comprehensive pdf documents analysis. The purpose of that subset differs in that regard, as focus has been done on making the dataset machine learning-ready for vision-language models.
An example page of one pdf document, with added bounding… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/pdfa-eng-wds.nara_revolutionary_war_pension_files_PDFs
Dataset Card for American Revolutionary War Pension Files - File-Level
Dataset Summary
A dataset derived from the National Archives and Records Administration (NARA) series Case Files of Pension and Bounty-Land Warrant Applications Based on American Revolutionary War Service (NARA Catalog Series, NAID 300022). This dataset provides a file-level representation of Revolutionary War pension records, aggregating individual page records into complete pension files… See the full description on the dataset page: https://huggingface.co/datasets/RevolutionCrossroads/nara_revolutionary_war_pension_files_PDFs.MINT-1T-PDF-CC-2023-23
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.MINT-1T-PDF-CC-2024-10
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.MINT-1T-PDF-CC-2023-40
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-40.msf-pdfsphaseShift_shell_result_pdf
Phase Resonance / IRS-DCE
Topological Dynamics & Artificial Cognitive Physics
Open Structural Record of Basis-Relative Reorganization in Transformer Representation Space
All pdf Creative Creative Commons Attribution No Derivatives 4.0 International
This repository provides comprehensive PDF research materials and Python scripts and something for mathematical proofs in the field of AI.
📢 Achievement: total 51k+ Downloads!… See the full description on the dataset page: https://huggingface.co/datasets/meta13sphere/phaseShift_shell_result_pdf.MINT-1T-PDF-CC-2023-06
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-06.MINT-1T-PDF-CC-2024-18
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-18.MINT-1T-PDF-CC-2023-14
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.pdfa-eng-wds
Dataset Card for PDF Association dataset (PDFA)
Dataset Summary
PDFA dataset is a document dataset filtered from the SafeDocs corpus, aka CC-MAIN-2021-31-PDF-UNTRUNCATED. The original purpose of that corpus is for comprehensive pdf documents analysis. The purpose of that subset differs in that regard, as focus has been done on making the dataset machine learning-ready for vision-language models.
An example page of one pdf document, with added bounding boxes… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/pdfa-eng-wds.SCPWiki-Cleaned-PDF-Archivesuparxive_boxed_pdfpdfQA-Benchmark
pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs
pdfQA is a structured benchmark collection for document-level question answering and PDF understanding research.
The dataset is organized to support:
Raw document processing research
Structured extraction pipelines
Retrieval-augmented QA
End-to-end document reasoning systems
It preserves original documents alongside structured derivatives to enable reproducible evaluation across preprocessing strategies.… See the full description on the dataset page: https://huggingface.co/datasets/pdfqa/pdfQA-Benchmark.20000-yingshi-pdfMINT-1T-PDF-CC-2023-50
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-50.bai-pdfslibretexts_pdfsdolmino_olmocr_pdfsgovdocs1-pdf-source
govdocs1: source PDF files
[!NOTE]
Converted versions of other document types (word, txt, etc) are available in this repo
This is ~220,000 open-access PDF documents (about 6.6M pages) from the dataset govdocs1. It wants to be OCR'd.
Uploaded as tar file pieces of ~10 GiB each due to size/file count limits with an index.csv covering details
5,000 randomly sampled PDFs are available unarchived in sample/. Hugging Face supports previewing these in-browser, for example this one… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-pdf-source.output_pdf_lmdbshorouk-pdf-archive
Shorouk PDF archive
Original newspaper PDFs from Al Shorouk's public archive. Filenames use YYYY-MM-DD.pdf.
As verified in this update, the dataset contains 6,276 PDFs, dated 2009-02-01 through 2026-09-08. The September 8 edition was already available from the publisher when checked on September 7 in New York.
Coverage and remaining gaps
This update added 518 issues and filled all calendar gaps from January 1, 2025 through September 8, 2026. 153 earlier dates… See the full description on the dataset page: https://huggingface.co/datasets/eltokh7/shorouk-pdf-archive.supreme_court_opinions_corpus_pdfwebAug24marianne_pdf_7marianne_pdf_9
