BEE-spoke-data/govdocs1-pdf-source
govdocs1: source PDF files [!NOTE] Converted versions of other document types (word, txt, etc) are available in this repo This is ~220,000 open-access PDF documents (about 6.6M pages) from the dataset govdocs1. It wants to be OCR'd. Uploaded as tar file pieces of ~10 GiB each due to size/file count limits with an index.csv covering details 5,000 randomly sampled PDFs are available unarchived in sample/. Hugging Face supports previewing these in-browser, for example this one… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-pdf-source.
govdocs1: source PDF files
[!NOTE] Converted versions of other document types (word, txt, etc) are available in this repo
This is ~220,000 open-access PDF documents (about 6.6M pages) from the dataset govdocs1. It wants to be OCR'd.
- Uploaded as
tarfile pieces of ~10 GiB each due to size/file count limits with an index.csv covering details - 5,000 randomly sampled PDFs are available unarchived in
sample/. Hugging Face supports previewing these in-browser, for example this one
Recovering the data
Download the data/ directory (with huggingface-cli download or similar) extract the tar pieces:
cat data_pdfs_part.tar.* | tar -xf - && rm data_pdfs_part.tar.*processing details
duplicates
exact duplicate PDFs were removed with jdupes. See the log file for details.
By the numbers
Based on the index.csv
Dataset Overview
Document Structure
Page Count Distribution
File Size Distribution
Metadata Completeness Crisis
Title Quality Breakdown
Top Authors
Top Subjects
Processing Errors
Temporal Coverage
Critical Assessment
[!NOTE] Generated by Claude Sonnet-4, unsolicited (as always)
Data Quality Issues
Key Insights
Document Profile: Typical government PDF = 10 pages, 0.15 MB, metadata-poor
Fatal Flaw: This dataset has excellent technical extraction (99.96% success) but catastrophic intellectual organization. You're essentially working with 230K unlabeled documents.
Bottom Line: The structural data is solid, but without subject classification for 79% of documents, this is an unindexed digital landfill masquerading as an archive.
