datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Epstein-Files
The Epstein Files Dataset
Dataset Description
Dataset Summary
This dataset comprises a curated collection of publicly available documents and related materials concerning Jeffrey Epstein. It includes unsealed court filings, FBI reports, DOJ publications, and other official investigative records. These files have been aggregated from reputable public sources, such as the U.S. Department of Justice Epstein Library, House Oversight Committee releases, and… See the full description on the dataset page: https://huggingface.co/datasets/Nikity/Epstein-Files.epstein-data
Epstein DOJ Document Archive v2
1.42 million OCR'd documents from the Department of Justice Jeffrey Epstein document release, with structured entity extraction, vector embeddings, financial transactions, communication records, and a forensic audit trail.
Frontend: epstein.academy
What's New in v2
10.6M entities (up from 8.5M) — expanded NER extraction
2.1M chunk embeddings (up from 1.96M) — more documents embedded
49,770 financial transactions — credit card and bank… See the full description on the dataset page: https://huggingface.co/datasets/kabasshouse/epstein-data.Epstein-Files
The Epstein Files Dataset
Dataset Description
Dataset Summary
This dataset comprises a curated collection of publicly available documents and related materials concerning Jeffrey Epstein. It includes unsealed court filings, FBI reports, DOJ publications, and other official investigative records. These files have been aggregated from reputable public sources, such as the U.S. Department of Justice Epstein Library, House Oversight Committee releases, and… See the full description on the dataset page: https://huggingface.co/datasets/aurora2424/Epstein-Files.Epstein-Files
The Epstein Files Dataset
Dataset Description
Dataset Summary
This dataset comprises a curated collection of publicly available documents and related materials concerning Jeffrey Epstein. It includes unsealed court filings, FBI reports, DOJ publications, and other official investigative records. These files have been aggregated from reputable public sources, such as the U.S. Department of Justice Epstein Library, House Oversight Committee releases, and… See the full description on the dataset page: https://huggingface.co/datasets/ramvorg/Epstein-Files.Epstein-Files
The Epstein Files Dataset
Dataset Description
Dataset Summary
This dataset comprises a curated collection of publicly available documents and related materials concerning Jeffrey Epstein. It includes unsealed court filings, FBI reports, DOJ publications, and other official investigative records. These files have been aggregated from reputable public sources, such as the U.S. Department of Justice Epstein Library, House Oversight Committee releases, and… See the full description on the dataset page: https://huggingface.co/datasets/mindhug/Epstein-Files.Epstein-Files
The Epstein Files Dataset
Dataset Description
Dataset Summary
This dataset comprises a curated collection of publicly available documents and related materials concerning Jeffrey Epstein. It includes unsealed court filings, FBI reports, DOJ publications, and other official investigative records. These files have been aggregated from reputable public sources, such as the U.S. Department of Justice Epstein Library, House Oversight Committee releases, and… See the full description on the dataset page: https://huggingface.co/datasets/ohmygaugh/Epstein-Files.FULL_EPSTEIN_INDEX
FULL_EPSTEIN_INDEX
CONTENT WARNING: This repository contains graphic and highly sensitive material regarding sexual abuse, exploitation, trafficking, and violence. It also contains unverified allegations and raw witness statements. User discretion is strongly advised.
Overview
Note. There is ALOT of data. OCR made mistakes scanning the files. So that being said, there is a lot of noise in the dataset, whether it be from OCR taking words out of normal… See the full description on the dataset page: https://huggingface.co/datasets/theelderemo/FULL_EPSTEIN_INDEX.Epstein-Files
The Epstein Files Dataset
Dataset Description
Dataset Summary
This dataset comprises a curated collection of publicly available documents and related materials concerning Jeffrey Epstein. It includes unsealed court filings, FBI reports, DOJ publications, and other official investigative records. These files have been aggregated from reputable public sources, such as the U.S. Department of Justice Epstein Library, House Oversight Committee releases, and… See the full description on the dataset page: https://huggingface.co/datasets/AfricanKillshot/Epstein-Files.epstein-files-ocr-complete
Epstein Files — Complete OCR Dataset
This is a comprehensive, structured publication of the Epstein Files OCR dataset, significantly expanding upon the earlier Datasets 1-8 release.
Dataset Summary
This dataset contains page-level OCR output compiled from an extensive release of documents related to Jeffrey Epstein / the Epstein case.
Each row in this dataset represents one scanned PDF document from the original release using a proprietary automated OCR pipeline… See the full description on the dataset page: https://huggingface.co/datasets/ishumilin/epstein-files-ocr-complete.epstein-emails
Epstein Email Messages Dataset
Dataset Summary
This dataset contains 4,272 individual email messages extracted directly from screenshot images using advanced Vision LLM technology. Unlike other datasets that work with only pre-OCR'd text, this dataset also used the original JPG screenshot images from the U.S. House Oversight Committee release using Qwen 2.5 VL 72B vision model to extract structured email data with high accuracy.
The dataset powers a live Progressive Web… See the full description on the dataset page: https://huggingface.co/datasets/to-be/epstein-emails.Epstein-Files
The Epstein Files Dataset
Dataset Description
Dataset Summary
This dataset comprises a curated collection of publicly available documents and related materials concerning Jeffrey Epstein. It includes unsealed court filings, FBI reports, DOJ publications, and other official investigative records. These files have been aggregated from reputable public sources, such as the U.S. Department of Justice Epstein Library, House Oversight Committee releases, and… See the full description on the dataset page: https://huggingface.co/datasets/post-train/Epstein-Files.epstein-emails
Epstein Email Threads Dataset
Dataset Summary
This dataset contains 5,082 parsed email threads extracted from OCR'd documents released by the U.S. House Oversight Committee. The emails have been processed using large language models to extract structured information including senders, recipients, timestamps, subjects, and message bodies, with OCR errors corrected and footers removed.
Dataset Description
Overview
This is a structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/notesbymuneeb/epstein-emails.epstein-files-20k
Disclaimer
This dataset is a reupload of a previously circulating public dataset. The contents may include unverified, incomplete, disputed, or inaccurate information and should not be interpreted as factual, authoritative, or as proof of guilt for any individual.
This dataset is provided strictly for research, archival, and educational purposes, such as analysis of information propagation, data preservation, or media studies.
No claims are made regarding the accuracy, authenticity… See the full description on the dataset page: https://huggingface.co/datasets/teyler/epstein-files-20k.epstein-files-ocr-complete
Epstein Files — Complete OCR Dataset
This is a comprehensive, structured publication of the Epstein Files OCR dataset, significantly expanding upon the earlier Datasets 1-8 release.
Dataset Summary
This dataset contains page-level OCR output compiled from an extensive release of documents related to Jeffrey Epstein / the Epstein case.
Each row in this dataset represents one scanned PDF document from the original release using a proprietary automated OCR pipeline… See the full description on the dataset page: https://huggingface.co/datasets/genevera/epstein-files-ocr-complete.epstein-emails
Epstein Email Threads Dataset
Dataset Summary
This dataset contains 5,082 parsed email threads extracted from OCR'd documents released by the U.S. House Oversight Committee. The emails have been processed using large language models to extract structured information including senders, recipients, timestamps, subjects, and message bodies, with OCR errors corrected and footers removed.
Dataset Description
Overview
This is a structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/Hannah2704/epstein-emails.epstein-files-ocr-complete
Epstein Files — Complete OCR Dataset
This is a comprehensive, structured publication of the Epstein Files OCR dataset, significantly expanding upon the earlier Datasets 1-8 release.
Dataset Summary
This dataset contains page-level OCR output compiled from an extensive release of documents related to Jeffrey Epstein / the Epstein case.
Each row in this dataset represents one scanned PDF document from the original release using a proprietary automated OCR pipeline… See the full description on the dataset page: https://huggingface.co/datasets/jpeglle/epstein-files-ocr-complete.amazon_polarity_10_pct
Amazon Polarity 10pct
This is a direct subset of the original Amazon Polarity dataset, downsampled 10pct with a random shuffle
Dataset Summary
For quicker testing on Amazon Polarity. See https://huggingface.co/datasets/amazon_polarity for details and attributions
Source Data
from datasets import ClassLabel, Dataset, DatasetDict, load_dataset
ds_full = load_dataset("amazon_polarity", streaming=True)
ds_train_10_pct =… See the full description on the dataset page: https://huggingface.co/datasets/ben-epstein/amazon_polarity_10_pct.epstein-files-20k
Disclaimer
This dataset is a reupload of a previously circulating public dataset. The contents may include unverified, incomplete, disputed, or inaccurate information and should not be interpreted as factual, authoritative, or as proof of guilt for any individual.
This dataset is provided strictly for research, archival, and educational purposes, such as analysis of information propagation, data preservation, or media studies.
No claims are made regarding the accuracy, authenticity… See the full description on the dataset page: https://huggingface.co/datasets/rbinrs/epstein-files-20k.epstein-files-nov-2025
Search the documents online - Hosted on Hugging Face. If the site is down, please create an issue to let me know.
NEW DATASET WITH LATEST RELEASES AND COMBINED FILES
U.S. House Oversight Epstein Estate Documents
Overview
On November 12, 2025, the U.S. House Oversight Committee released over 20,000 pages of documents from the Epstein estate. While intended to serve the public interest, these records remain largely inaccessible as they are scattered across nested… See the full description on the dataset page: https://huggingface.co/datasets/theelderemo/epstein-files-nov-2025.Jeffrey-Epstein-Emails-From-Epstein-Files
Jeffrey Epstein Emails from Epstein Files
This dataset contains 7,380 emails with 14,835 messages scraped from jmail.world.
Dataset Description
This dataset provides email correspondence from the Jeffrey Epstein Files, scraped directly from the jmail.world website.
Data Source
All emails were scraped from the jmail.world website.
Fields
Field
Type
Description
doc_id
string
Internal document identifier
subject
string
Email subject line… See the full description on the dataset page: https://huggingface.co/datasets/KillerShoaib/Jeffrey-Epstein-Emails-From-Epstein-Files.epstein-files-20k
Disclaimer
This dataset is a reupload of a previously circulating public dataset. The contents may include unverified, incomplete, disputed, or inaccurate information and should not be interpreted as factual, authoritative, or as proof of guilt for any individual.
This dataset is provided strictly for research, archival, and educational purposes, such as analysis of information propagation, data preservation, or media studies.
No claims are made regarding the accuracy, authenticity… See the full description on the dataset page: https://huggingface.co/datasets/aurora2424/epstein-files-20k.epstein-files-nov11-25-house-post-ocr-embeddings
Epstein Files Document Embeddings
Text embeddings generated from the House Oversight Committee's Epstein document release.
Source Dataset
This dataset is derived from: tensonaut/EPSTEIN_FILES_20K
The source dataset contains OCR'd text from the original House Oversight Committee PDF release.
Dataset Structure
Field
Type
Description
source_file
string
Source document filename
chunk_index
int
Position of chunk within document
text
string
Original… See the full description on the dataset page: https://huggingface.co/datasets/svetfm/epstein-files-nov11-25-house-post-ocr-embeddings.epstein-emails
Epstein Email Threads Dataset
Dataset Summary
This dataset contains 5,082 parsed email threads extracted from OCR'd documents released by the U.S. House Oversight Committee. The emails have been processed using large language models to extract structured information including senders, recipients, timestamps, subjects, and message bodies, with OCR errors corrected and footers removed.
Dataset Description
Overview
This is a structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/pupepps/epstein-emails.epstein-data-chatepstein-imagesepstein-emailFULL_EPSTEIN_INDEX
FULL_EPSTEIN_INDEX
CONTENT WARNING: This repository contains graphic and highly sensitive material regarding sexual abuse, exploitation, trafficking, and violence. It also contains unverified allegations and raw witness statements. User discretion is strongly advised.
Overview
Note. There is ALOT of data. OCR made mistakes scanning the files. So that being said, there is a lot of noise in the dataset, whether it be from OCR taking words out of normal… See the full description on the dataset page: https://huggingface.co/datasets/genevera/FULL_EPSTEIN_INDEX.epstein-images-croppedepstein-dojvectorized_epstein_20k
Epstein Vectorized Database
This repository contains a vectorized derivative of publicly released U.S. House Oversight Committee Epstein estate documents. The database provides efficient semantic search and retrieval-augmented generation (RAG) capabilities via FAISS embeddings and metadata.
Database Contents
epstein_index.faiss: FAISS vector index (384-dim, normalized embeddings via all-MiniLM-L6-v2, IndexFlatIP for cosine similarity)
epstein_metadata.parquet: Metadata… See the full description on the dataset page: https://huggingface.co/datasets/nglif/vectorized_epstein_20k.
