datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
epstein-files-ocr-datasets-1-8-early-release
Epstein Files OCR — Datasets 1–8 (Early Release)
ARCHIVE NOTICE
This dataset is no longer maintained. Please refer to the Epstein Files — Complete OCR Dataset.
Dataset Summary
This dataset contains page-level OCR output (as Markdown) from a public release of documents related to Jeffrey Epstein / the Epstein case.
Each Markdown file represents one scanned page converted to text using an automated OCR pipeline. The dataset is designed for:
Question answering
Information… See the full description on the dataset page: https://huggingface.co/datasets/ishumilin/epstein-files-ocr-datasets-1-8-early-release.epstein-files-ocr-datasets-1-8-early-release
Epstein Files OCR — Datasets 1–8 (Early Release)
Work in Progress (WIP)
This is an early publication. We are actively working on improving OCR quality and expanding coverage.
Dataset Summary
This dataset contains page-level OCR output (as Markdown) from a public release of documents related to Jeffrey Epstein / the Epstein case.
Each Markdown file represents one scanned page converted to text using an automated OCR pipeline. The dataset is designed for:
Question… See the full description on the dataset page: https://huggingface.co/datasets/aurora2424/epstein-files-ocr-datasets-1-8-early-release.epstein-files-ocr-complete
Epstein Files — Complete OCR Dataset
This is a comprehensive, structured publication of the Epstein Files OCR dataset, significantly expanding upon the earlier Datasets 1-8 release.
Dataset Summary
This dataset contains page-level OCR output compiled from an extensive release of documents related to Jeffrey Epstein / the Epstein case.
Each row in this dataset represents one scanned PDF document from the original release using a proprietary automated OCR pipeline… See the full description on the dataset page: https://huggingface.co/datasets/ishumilin/epstein-files-ocr-complete.epstein-files-ocr-complete
Epstein Files — Complete OCR Dataset
This is a comprehensive, structured publication of the Epstein Files OCR dataset, significantly expanding upon the earlier Datasets 1-8 release.
Dataset Summary
This dataset contains page-level OCR output compiled from an extensive release of documents related to Jeffrey Epstein / the Epstein case.
Each row in this dataset represents one scanned PDF document from the original release using a proprietary automated OCR pipeline… See the full description on the dataset page: https://huggingface.co/datasets/genevera/epstein-files-ocr-complete.epstein-files-ocr-complete
Epstein Files — Complete OCR Dataset
This is a comprehensive, structured publication of the Epstein Files OCR dataset, significantly expanding upon the earlier Datasets 1-8 release.
Dataset Summary
This dataset contains page-level OCR output compiled from an extensive release of documents related to Jeffrey Epstein / the Epstein case.
Each row in this dataset represents one scanned PDF document from the original release using a proprietary automated OCR pipeline… See the full description on the dataset page: https://huggingface.co/datasets/jpeglle/epstein-files-ocr-complete.epstein-files-nov11-25-house-post-ocr-embeddings
Epstein Files Document Embeddings
Text embeddings generated from the House Oversight Committee's Epstein document release.
Source Dataset
This dataset is derived from: tensonaut/EPSTEIN_FILES_20K
The source dataset contains OCR'd text from the original House Oversight Committee PDF release.
Dataset Structure
Field
Type
Description
source_file
string
Source document filename
chunk_index
int
Position of chunk within document
text
string
Original… See the full description on the dataset page: https://huggingface.co/datasets/svetfm/epstein-files-nov11-25-house-post-ocr-embeddings.epstein-files-nov11-25-house-post-ocr-embeddings
Epstein Files Document Embeddings
Text embeddings generated from the House Oversight Committee's Epstein document release.
Source Dataset
This dataset is derived from: tensonaut/EPSTEIN_FILES_20K
The source dataset contains OCR'd text from the original House Oversight Committee PDF release.
Dataset Structure
Field
Type
Description
source_file
string
Source document filename
chunk_index
int
Position of chunk within document
text
string
Original… See the full description on the dataset page: https://huggingface.co/datasets/TinaWild09/epstein-files-nov11-25-house-post-ocr-embeddings.
