datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medpmc-11m-dataset_jun24_baseline
MedPMC WebDataset
MedPMC is a large-scale medical image-text dataset curated from articles in the PubMed Central (PMC) collection. This release contains approximately 11 million image-text pairs collected from the June 2024 PMC baseline. MedPMC is an ongoing effort, and future releases will continue to expand the dataset with newly published literature, improved annotations, and additional resources.
This dataset is presented in the paper MedPMC: A Systematic Framework for… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/medpmc-11m-dataset_jun24_baseline.olmoearth_pretrain_datasetThis is the pre-training dataset for training the OlmoEarth pre-trained remote sensing foundation models.
Documentation is on GitHub at https://github.com/allenai/olmoearth_pretrain/blob/main/docs/Pretraining-Dataset.md
The dataset is released under CC BY 4.0. It includes data from the following sources:
Sentinel-2 L2A imagery from the European Space Agency, available under the Copernicus Sentinel Data and Service Legal Notice
Sentinel-1 GRD IW vv+vh imagery from the European Space Agency… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmoearth_pretrain_dataset.X2Edit-Dataset
X2Edit
Introduction
X2Edit Dataset is a comprehensive image editing dataset that covers 14 diverse editing tasks and exhibits substantial advantages over existing open-source datasets including AnyEdit, HQ-Edit, UltraEdit, SEED-Data-Edit, ImgEdit and OmniEdit.
For the relevant data construction scripts, model training and inference scripts, please refer to X2Edit.
News
2025/09/16: We are about to release a dataset constructed by Qwen-Image and… See the full description on the dataset page: https://huggingface.co/datasets/OPPOer/X2Edit-Dataset.findtextCenterNet_datasetdescribe-anything-dataset
Describe Anything: Detailed Localized Image and Video Captioning
NVIDIA, UC Berkeley, UCSF
Long Lian, Yifan Ding, Yunhao Ge, Sifei Liu, Hanzi Mao, Boyi Li, Marco Pavone, Ming-Yu Liu, Trevor Darrell, Adam Yala, Yin Cui
[Paper] | [Code] | [Project Page] | [Video] | [HuggingFace Demo] | [Model/Benchmark/Datasets] | [Citation]
Dataset Card for Describe Anything Datasets
Datasets used in the training of describe anything models (DAM).
The datasets are in tar files. These… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/describe-anything-dataset.STRIDE-QA-Dataset
STRIDE-QA Dataset
📦 Dataset
STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks.
Category
Description
Object-centric Spatial QA
Spatial relations between two… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Dataset.Harmonizer-Dataset
HARMONIZER DATASET
Dataset Description
Training dataset for DiffusionHarmonizer: a generative AI model for image and video enhancement bridging neural reconstruction and photorealistic simulation .
Model checkpoints: https://huggingface.co/nvidia/Harmonizer/Training code: https://github.com/NVIDIA/harmonizer/
The dataset was curated to support the following functions of the model:
3D reconstruction artifact removal
Harmonization of inserted objects to blend… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Harmonizer-Dataset.JANuS_datasetThis repository hosts the JANuS (Joint Annotations and Names) dataset introduced in the 2023 paper Distributionally Robust Classification on a Data Budget.
As of this writing, ours is the only public dataset which is both fully annotated with ground-truth labels and fully captioned with web-scraped captions.
It is designed to be used for controlled experiments with vision-language models.
What is in JANuS?
JANuS provides metadata and image links for four new training datasets; all… See the full description on the dataset page: https://huggingface.co/datasets/penfever/JANuS_dataset.plotqa-dataset
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Dodon/plotqa-dataset.PIG-Nav-Dataset-PretrainThe pretraining dataset for our paper PIG-Nav: Key Insights for Pretrained Image-Goal Navigation Models.
Description of the dataset:
The pretraining dataset include GoStanford, RECON, CoryHall, Berkeley DeepDrive, SCAND, TartanDrive, and SACSoN.
Please follow our github repo https://github.com/zpschang/PIG-Nav for detailed use.
RORem_datasetHO-Cap-Datasetwaifu-preprocessed-datasetforbin_dataset
Forbin Dataset: A collection of historical photographs with archival metadata
This repository hosts the Forbin Dataset, a large-scale collection of historical photographs taken or collected by Victor Forbin (1868–1947).
This HuggingFace dataset version provides:
COCO-style annotations (segmentation polygons)
Archival metadata (Box ID, description, notes, dates when available)
A lightweight explorer interface (HTML/JS) to preview images and annotations:… See the full description on the dataset page: https://huggingface.co/datasets/mchelali/forbin_dataset.The_Million_Song_Datasetyuxuan_good_dataset_dtUnified_Road_Defect_Dataset
Unified Road Defect Dataset
A merged, YOLO-format road-defect detection dataset that combines RDD-2022
(primary, ground-level, 6 countries) with two supplementary aerial/drone
datasets — UAV-PDD2023 (China) and RoadDamageVision (China + Spain) —
into a single 4-class CRDDC schema.
This is a derived dataset. It re-packages and re-labels images from three
independently published sources. All credit for the underlying images and
original annotations belongs to their respective… See the full description on the dataset page: https://huggingface.co/datasets/TamAko783/Unified_Road_Defect_Dataset.drunet_datasetvigorl_datasets
ViGoRL Datasets
This repository contains the official datasets associated with the paper "Grounded Reinforcement Learning for Visual Reasoning (ViGoRL)", by Gabriel Sarch, Snigdha Saha, Naitik Khandelwal, Ayush Jain, Michael J. Tarr, Aviral Kumar, and Katerina Fragkiadaki.
Dataset Overview
These datasets are designed for training and evaluating visually grounded vision-language models (VLMs).
Datasets are organized by the visual reasoning tasks described in the ViGoRL… See the full description on the dataset page: https://huggingface.co/datasets/gsarch/vigorl_datasets.yuxuan_good_dataset_sttimage-manipulation-dataset-compilationjapanese-str-dataset-v1
STR Dataset
Japanese STR (Scene Text Recognition) dataset in WebDataset format.
This dataset is composed of:
Images of Japanese named entities (full names and their affiliations)
Images of sentences retrieved from Aozora Bunko (青空文庫)
and their corresponding ground truth texts.
All images are synthesized using TRDG.
Dataset Structure
Split
Samples
Shards
train
10,000,000
1000
valid
50,000
5
test
50,000
5
Total
10,100,000
1010
Usage
import… See the full description on the dataset page: https://huggingface.co/datasets/nagohachi/japanese-str-dataset-v1.Lowlight-Smartphone-Dataset
[WACV'26] Low-light Smartphone Dataset (LSD)
This is the official dataset proposed in our paper titled "Illuminating Darkness: Learning to Enhance Low-light Images In-the-Wild"
📄 Paper: arXiv💻 Code: GitHub - LSD-TFFormer
Overview
We introduce LSD, the largest in-the-wild Single-Shot Low-Light Image Enhancement (SLLIE) dataset to date.
Dataset Structure
This repository contains the following training data files:
patch_DLL_gtPatch.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/ARM4588/Lowlight-Smartphone-Dataset.CCPD-DatasetTHIS DATABASE WAS DEVELOPED BY THE CHINESE UNIVERSITY OF SCIENCE AND TECHNOLOGY, MY ONLY ROLE IN THIS STAGE HAS ONLY BEEN UPLOADING IT TO HUGGING FACE. ALL CREDIT IS DUE TO THE UNIVERSITY OF SCIENCE AND TECHNOLOGY
DF_DiFF_FAS_dataset_in_FSFM_FSVFM
FSFM / FS-VFM Downstream Datasets
Processed downstream fine-tuning datasets for cross-dataset deepfake detection, cross-domain face anti-spoofing, and unseen diffusion-generated face detection used with FSFM and FS-VFM.
This archive is intended for research use with the FSFM/FS-VFM release scripts.
curvton_dataset
Curvton Dataset
Dataset Summary
Curvton is a synthetic virtual try-on dataset hosted on Hugging Face as six tar archives:
easy_female.tar
easy_male.tar
medium_female.tar
medium_male.tar
hard_female.tar
hard_male.tar
Each archive stores JPEG files in the following internal structure:
<difficulty>/<gender>/
cloth_image/
initial_person_image/
tryon_image/
The current published repository is archive-oriented rather than parquet- or CSV-oriented. This means the… See the full description on the dataset page: https://huggingface.co/datasets/curvton/curvton_dataset.SD_inpaint_datasetgrounding-YT-dataset
Grounding YouTube Dataset
What, when, and where? -- Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions
arxiv
This dataset is packed in WebDataset format.
The dataset is present in three styles:
Untrimmed videos + annotations within the entire video
Action clips extracted from the videos + annotations in each clip
Action frames extracted from the videos + annotation of the frame
Example usage for clips:… See the full description on the dataset page: https://huggingface.co/datasets/CVML-TueAI/grounding-YT-dataset.flickr30k-datasetyuxuan_dataset_dt
