datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
foia-reading-room-documents
Foia Reading Room Documents
Documents from federal FOIA reading rooms and Inspector General report libraries: audits, inspections, investigative summaries and records released under the Freedom of Information Act.
Every document here was published by a US federal agency and is a work of the
United States government. Nothing has been altered: files are byte-identical to
what the agency posted, and the checksum in metadata.parquet is of the
original bytes.
Why this… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/foia-reading-room-documents.mkultra-foia-mineru-archive
MKULTRA FOIA MinerU Archive
The MKULTRA FOIA MinerU Archive is a public-interest research dataset
containing 21,237 pages of declassified MKULTRA and related-program
records.
The repository contains two separately preserved collections:
Collection
Pages
Source
Page export
Document export
Original PEERS FOIA collection
16,383
TIFF files obtained by PEERS through FOIA and converted to PNG
pages
documents
Supplemental authenticated declassified collection
4,854… See the full description on the dataset page: https://huggingface.co/datasets/Thorismund/mkultra-foia-mineru-archive.mkultra-foia-mineru-archive
MKULTRA FOIA MinerU Archive
The MKULTRA FOIA MinerU Archive is a public-interest research dataset
containing 21,237 pages of declassified MKULTRA and related-program
records.
The repository contains two separately preserved collections:
Collection
Pages
Source
Page export
Document export
Original PEERS FOIA collection
16,383
TIFF files obtained by PEERS through FOIA and converted to PNG
pages
documents
Supplemental authenticated declassified collection
4,854… See the full description on the dataset page: https://huggingface.co/datasets/peers-ai/mkultra-foia-mineru-archive.FOIA-2foiarchive
Dataset Summary
Our multidisciplinary team of researchers has gathered nearly 5 million documents, comprising over 18 million pages, to create the Freedom of Information Archive (FOIArchive), the world's largest database of declassified government records.
The FOIA Archive data contains the full text and selected metadata from History Lab's collection of declassified government documents.
Dataset structure
The data are in JSONL format with 9 different fields.… See the full description on the dataset page: https://huggingface.co/datasets/HistoryLab/foiarchive.
