epstein
Datasets
All datasets matching “epstein”epstein-files-ocr-datasets-1-8-early-release
Epstein Files OCR — Datasets 1–8 (Early Release)
ARCHIVE NOTICE
This dataset is no longer maintained. Please refer to the Epstein Files — Complete OCR Dataset.
Dataset Summary
This dataset contains page-level OCR output (as Markdown) from a public release of documents related to Jeffrey Epstein / the Epstein case.
Each Markdown file represents one scanned page converted to text using an automated OCR pipeline. The dataset is designed for:
Question answering
Information… See the full description on the dataset page: https://huggingface.co/datasets/ishumilin/epstein-files-ocr-datasets-1-8-early-release.epstein-fbi-files
FBI Epstein Files - Embeddings Dataset
Document embeddings and OCR text from the FBI's release of Jeffrey Epstein-related files.
Dataset Structure
embeddings/
all_embeddings.jsonl # 236K chunks with 768-dim embeddings
ocr/
all_ocr.jsonl # Full OCR text for each document
Embedding Format
Each line in all_embeddings.jsonl is a JSON object:
{
"id": "uuid",
"bates_number": "EFTA00000001",
"bates_range": "EFTA00000001-EFTA00000001"… See the full description on the dataset page: https://huggingface.co/datasets/svetfm/epstein-fbi-files.Epstein-Files
The Epstein Files Dataset
Dataset Description
Dataset Summary
This dataset comprises a curated collection of publicly available documents and related materials concerning Jeffrey Epstein. It includes unsealed court filings, FBI reports, DOJ publications, and other official investigative records. These files have been aggregated from reputable public sources, such as the U.S. Department of Justice Epstein Library, House Oversight Committee releases, and… See the full description on the dataset page: https://huggingface.co/datasets/Nikity/Epstein-Files.epstein-index
Epstein Files Corpus — OCR'd & Cleaned
Machine-readable text of the publicly released Jeffrey Epstein–related document
collections: the DOJ "EFTA" disclosures (Data Sets 1–12), House Oversight
releases, court records (Maxwell, Doe v. Epstein/Indyke, Florida v. Epstein,
USVI v. JPMorgan, and others), and FOIA productions (FBI, BOP, CBP, Florida).
All content is public-record material released by the U.S. Department of
Justice, federal/state courts, and FOIA respondents.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/robbd/epstein-index.epstein-data
Epstein DOJ Document Archive v2
1.42 million OCR'd documents from the Department of Justice Jeffrey Epstein document release, with structured entity extraction, vector embeddings, financial transactions, communication records, and a forensic audit trail.
Frontend: epstein.academy
What's New in v2
10.6M entities (up from 8.5M) — expanded NER extraction
2.1M chunk embeddings (up from 1.96M) — more documents embedded
49,770 financial transactions — credit card and bank… See the full description on the dataset page: https://huggingface.co/datasets/kabasshouse/epstein-data.epstein-files-ocr-datasets-1-8-early-release
Epstein Files OCR — Datasets 1–8 (Early Release)
Work in Progress (WIP)
This is an early publication. We are actively working on improving OCR quality and expanding coverage.
Dataset Summary
This dataset contains page-level OCR output (as Markdown) from a public release of documents related to Jeffrey Epstein / the Epstein case.
Each Markdown file represents one scanned page converted to text using an automated OCR pipeline. The dataset is designed for:
Question… See the full description on the dataset page: https://huggingface.co/datasets/aurora2424/epstein-files-ocr-datasets-1-8-early-release.
