datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ubuntu_osworld_file_cache
OSWorld File Cache
This repository serves as a file cache for the OSWorld project, providing reliable and fast access to evaluation files that were previously hosted on Google Drive.
Overview
OSWorld is a scalable, real computer environment for multimodal agents, supporting task setup, execution-based evaluation, and interactive learning across various operating systems and applications. This cache repository ensures that all evaluation files are consistently… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/ubuntu_osworld_file_cache.dataset_with_data_filesfilesnara_revolutionary_war_pension_files_PDFs
Dataset Card for American Revolutionary War Pension Files - File-Level
Dataset Summary
A dataset derived from the National Archives and Records Administration (NARA) series Case Files of Pension and Bounty-Land Warrant Applications Based on American Revolutionary War Service (NARA Catalog Series, NAID 300022). This dataset provides a file-level representation of Revolutionary War pension records, aggregating individual page records into complete pension files… See the full description on the dataset page: https://huggingface.co/datasets/RevolutionCrossroads/nara_revolutionary_war_pension_files_PDFs.gaia2_filesystem
GAIA2 Filesystem
This is a dataset containing files for the GAIA2 benchmark. You should not use this dataset on its own, but instead use the Meta Agents Research Environments framework to execute scenarios from that GAIA2 dataset.
Dataset Link
https://huggingface.co/datasets/meta-agents-research-environments/gaia2
Contact Details
Publishing POC: Meta AI Research Team
Affiliation: Meta Platforms, Inc.
Website:… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2_filesystem.eurospeech-raw-filescompressed_filescode_execution_filescode_python_filesfilecodeboxsif-files-for-swewindows_osworld_file_cache
OSWorld File Cache
This repository serves as a file cache for the OSWorld project, providing reliable and fast access to evaluation files that were previously hosted on Google Drive.
Overview
OSWorld is a scalable, real computer environment for multimodal agents, supporting task setup, execution-based evaluation, and interactive learning across various operating systems and applications. This cache repository ensures that all evaluation files are consistently accessible… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/windows_osworld_file_cache.custom_code_py_filesaudio-filesgated_dataset_with_data_filescsd_filesgta-data-files-universalcustom_code_execution_filesmy-filesepstein-files-ocr-datasets-1-8-early-release
Epstein Files OCR — Datasets 1–8 (Early Release)
ARCHIVE NOTICE
This dataset is no longer maintained. Please refer to the Epstein Files — Complete OCR Dataset.
Dataset Summary
This dataset contains page-level OCR output (as Markdown) from a public release of documents related to Jeffrey Epstein / the Epstein case.
Each Markdown file represents one scanned page converted to text using an automated OCR pipeline. The dataset is designed for:
Question answering
Information… See the full description on the dataset page: https://huggingface.co/datasets/ishumilin/epstein-files-ocr-datasets-1-8-early-release.Dyna_Repo_Inria_MDPosit_Files
DynaRepo MDPosit
This repository mirrors Molecular Dynamics (MD) project metadata and files laid out per accession as exposed by Dynarepo. It includes global and per-accession manifests, project-level JSON, and the actual files (structures, trajectories, derived PCA binaries, and analysis screenshots). Field names and shapes follow Dynarepo’s documented model.
Overview
How to consume
Stream-parse: Read the file line-by-line and json.loads each… See the full description on the dataset page: https://huggingface.co/datasets/DataQuests/Dyna_Repo_Inria_MDPosit_Files.nara_revolutionary_war_pension_files
Dataset Card for American Revolutionary War Pension Files
Dataset Summary
A dataset derived from the Case Files of Pension and Bounty-Land Warrant Applications Based on American Revolutionary War Service, ca. 1800–ca. 1912 (NARA Catalog Series, NAID 300022). This dataset includes page-level records with digitized images, original extracted text (by Family Search), AI-generated OCR, and human-created transcriptions where available. It offers a unique window into… See the full description on the dataset page: https://huggingface.co/datasets/RevolutionCrossroads/nara_revolutionary_war_pension_files.epstein-fbi-files
FBI Epstein Files - Embeddings Dataset
Document embeddings and OCR text from the FBI's release of Jeffrey Epstein-related files.
Dataset Structure
embeddings/
all_embeddings.jsonl # 236K chunks with 768-dim embeddings
ocr/
all_ocr.jsonl # Full OCR text for each document
Embedding Format
Each line in all_embeddings.jsonl is a JSON object:
{
"id": "uuid",
"bates_number": "EFTA00000001",
"bates_range": "EFTA00000001-EFTA00000001"… See the full description on the dataset page: https://huggingface.co/datasets/svetfm/epstein-fbi-files.audio-filestemp_filereamixed-project-files
reamixed_project_files
filegram-bench-datalabel-filesThis repository contains the mapping from integer id's to actual label names (in HuggingFace Transformers typically called id2label) for several datasets.
Current datasets include:
ImageNet-1k
ImageNet-22k (also called ImageNet-21k as there are 21,843 classes)
COCO detection 2017
COCO panoptic 2017
ADE20k (actually, the MIT Scene Parsing benchmark, which is a subset of ADE20k)
Cityscapes
VQAv2
Kinetics-700
RVL-CDIP
PASCAL VOC
Kinetics-400
...
You can read in a label file as follows (using… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/label-files.my-shared-filesfile.checkpoints
