datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
research-papers
research-papers Dataset
Overview
The Research Papers Dataset is a collection of academic research documents categorized by their primary research topic.
This dataset is designed for tasks such as model finetuning, document classification, optical character recognition (OCR) testing and multimodal document understanding (Feel free to use it however you see fit!).
Curated by: tegridy
Language: English
Format: PDF | MD
Repo Structure
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/tegridydev/research-papers.olmoearth-paper-embeddings
OlmoEarth — Foundation-Model Embeddings for Paper Table 2
This dataset contains pre-extracted embeddings from 26 Earth-observation
foundation models evaluated on the 24 downstream tasks that make up
Table 2 of the OlmoEarth paper:
OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation
AI2, 2025. arXiv:2511.13655.
For every supported (model, task) pair we ran the model's encoder over the
task's train / validation / test splits with the paper-best… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmoearth-paper-embeddings.3dvs2026_papers
Dataset Card for 3dvs2026_papers
This is a FiftyOne dataset with 176 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/3dvs2026_papers")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/3dvs2026_papers.research-papers
research-papers Dataset
Overview
The Research Papers Dataset is a collection of academic research documents in PDF format, categorized by their primary research topic. This dataset is designed for tasks such as document classification, optical character recognition (OCR) testing, and multimodal document understanding.
Curated by: tegridy
Language: English
Format: PDF
Repo Structure
The dataset contains PDF files and their associated topic labels.… See the full description on the dataset page: https://huggingface.co/datasets/rAJGAUTAMdsdsddsds12211212/research-papers.vecforge-paper-corpus
Note (rebuild in progress): figures are being re-extracted with a fixed extractor (cleaner crops). The image-preview config (Parquet with an inline column + difficulty/type/score labels) returns after re-classification. The config (paper metadata + links) is live now.
VecForge Paper Corpus
A pristine, deduplicated collection of 59,732 top-venue AI/ML/CV/NLP/Robotics papers (2020-2024) with
every captioned figure and full paper text, for research on figure understanding… See the full description on the dataset page: https://huggingface.co/datasets/debajyotidasgupta/vecforge-paper-corpus.research-papers
research-papers Dataset
Overview
The Research Papers Dataset is a collection of academic research documents in PDF format, categorized by their primary research topic. This dataset is designed for tasks such as document classification, optical character recognition (OCR) testing, and multimodal document understanding.
Curated by: tegridy
Language: English
Format: PDF
Repo Structure
The dataset contains PDF files and their associated topic labels.… See the full description on the dataset page: https://huggingface.co/datasets/itstheprakash/research-papers.HUGO-Bench-Paper-Reproducibility
HUGO-Bench Paper Reproducibility
Supplementary data and reproducibility materials for the paper:
Vision Transformers for Zero-Shot Clustering of Animal Images: A Comparative Benchmarking Study - https://arxiv.org/abs/2602.03894
Hugo Markoff, Stefan Hein Bengtson, Michael Ørsted
Aalborg University, Denmark
Dataset Description
This repository contains complete experimental results, pre-computed embeddings, and execution logs from our comprehensive benchmarking study… See the full description on the dataset page: https://huggingface.co/datasets/AI-EcoNet/HUGO-Bench-Paper-Reproducibility.3D_paper_mask_attack_dataset_for_Liveness
Liveness Detection Dataset: 3D Paper Mask Attacks
Paper mask attack dataset for training PAD and liveness detection models against low-cost 3D spoofing — 2,000+ videos on iOS and Android, ISO 30107-3 Level 1/2 attack category
What This Dataset Covers
2,000+ video recordings of 3D paper mask presentation attacks — masks with volumetric elements that simulate facial depth. Captured on iOS and Android devices, ~7 sec per video, with zoom-in/zoom-out phases for active… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/3D_paper_mask_attack_dataset_for_Liveness.hw1-paper-triage-multimodal
HW1 Multimodal Paper Triage
Purpose
This dataset supports a classroom exercise in assembling and augmenting multimodal data for a personalized research-paper triage system.
Composition and splits
The dataset begins with 100 original multimodal samples. The original samples were split before augmentation using a fixed random seed and stratification by the binary target.
train: 10,070 samples consisting of 70 training originals and 10,000 augmented… See the full description on the dataset page: https://huggingface.co/datasets/ishaanamahajan/hw1-paper-triage-multimodal.2d-paper-mask-face-anti-spoofing
Cut-Out Paper Mask Face Spoofing Dataset
3,000 videos of partial 2D paper mask attacks from 50 participants, recorded on Galaxy A54 and iPhone 14 Pro. Built for training and evaluating face anti-spoofing, liveness detection, and presentation attack detection systems.
Full dataset for commercial use — request a license at axonlab.ai
Why Cut-Out Masks Are a Harder Problem
Standard print attack and paper mask datasets use full-face photo printouts. Cut-out masks are… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/2d-paper-mask-face-anti-spoofing.DATA-DIFFUSION-BASED-PAPER
DATA-DIFFUSION-BASED-PAPER
Time-lapse fluorescence-microscopy data released with the DATA-DIFFUSION-BASED-PAPER project. The archive contains source image files and accompanying metadata used by the DINO benchmark and downstream analyses.
Contents
Experimental-condition directories containing TIFF frames and ZIP archives.
2,946 TIFF files and 78 ZIP archives.
3,024 metadata files.
Conditions include susceptible mono-culture + DTPA, DTPA controls, and… See the full description on the dataset page: https://huggingface.co/datasets/rubentium/DATA-DIFFUSION-BASED-PAPER.beyond_the_lab_neurips_paperHousehold-Paper-Products-Occlusion-Image-Dataset
Household Paper Products Occlusion Image Dataset
Currently, the retail e-commerce industry faces challenges in accurately detecting and categorizing household paper products due to varying packaging designs and occlusions in images. Existing datasets often lack diversity in terms of occlusion scenarios and packaging types, leading to reduced model performance in real-world applications. This dataset aims to address these issues by providing a comprehensive collection of images… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Household-Paper-Products-Occlusion-Image-Dataset.rock-paper-scissor-datasetrock-paper-scissor
