fbi
Datasets
All datasets matching “fbi”epstein-fbi-files
FBI Epstein Files - Embeddings Dataset
Document embeddings and OCR text from the FBI's release of Jeffrey Epstein-related files.
Dataset Structure
embeddings/
all_embeddings.jsonl # 236K chunks with 768-dim embeddings
ocr/
all_ocr.jsonl # Full OCR text for each document
Embedding Format
Each line in all_embeddings.jsonl is a JSON object:
{
"id": "uuid",
"bates_number": "EFTA00000001",
"bates_range": "EFTA00000001-EFTA00000001"… See the full description on the dataset page: https://huggingface.co/datasets/svetfm/epstein-fbi-files.FBI
Finding Blind Spots in Evaluator LLMs with Interpretable Checklists
We present FBI, our novel meta-evaluation framework designed to assess the robustness of evaluator LLMs across diverse tasks and evaluation strategies. Please refer to our paper for more details.
Code
The code to generate the perturbations and run evaluations are available on our github repository: ai4bharat/fbi
Tasks
We manually categorized each prompt into one of the 4 task… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/FBI.fb-imagesFBIS-22M
FBIS-22M
Field Boundary Instance Segmentation - 22M dataset (FBIS-22M) - large-scale, multi-resolution dataset comprising 672,909 high-resolution satellite image patches (0.25 m – 10 m) and 22,926,427 instance masks of individual fields.Note: We have excluded images from the Pleiades satellite mission due to third-party usage policies. This does not impact the dataset’s coverage or quality.
Dataset Structure
The dataset is split into multiple archive parts.… See the full description on the dataset page: https://huggingface.co/datasets/MykolaL/FBIS-22M.FBIS-73M
FBIS-73M
Field Boundary Instance Segmentation - 73M dataset (FBIS-73M) - large-scale, multi-resolution dataset comprising 1,478,096 high-resolution satellite image patches (0.25 m – 10 m) and 74,015,889 instance masks of individual fields.
Dataset Structure
The dataset has three splits: train, test, and a 100-country zero-shot test set. Each split is described by a list file, and the actual image/label data is distributed as independent zip archives.
File… See the full description on the dataset page: https://huggingface.co/datasets/MykolaL/FBIS-73M.kl3m-data-dotgov-www.fbi.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.fbi.gov.
