datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imagenet_1k_resized_256
Dataset Card for "imagenet_1k_resized_256"
Dataset summary
The same ImageNet dataset but all the smaller side resized to 256.
A lot of pretraining workflows contain resizing images to 256 and random cropping to 224x224, this is why 256 is chosen.
The resized dataset can also be downloaded much faster and consume less space than the original one.
See here for detailed readme.
Dataset Structure
Below is the example of one row of data. Note that the labels in… See the full description on the dataset page: https://huggingface.co/datasets/evanarlian/imagenet_1k_resized_256.resisc45
Description
RESISC45 dataset is a publicly available benchmark for Remote Sensing Image Scene Classification (RESISC), created by Northwestern Polytechnical University (NWPU). This dataset contains 31,500 images, covering 45 scene classes with 700 images in each class.
The dataset does not have any default splits. Train, validation, and test splits were based on these definitions here… See the full description on the dataset page: https://huggingface.co/datasets/timm/resisc45.Defactify_Image_Dataset
Defactify_Image_Dataset
This dataset is associated with the paper A Comprehensive Dataset for Human vs. AI Generated Image Detection.
📝 Dataset Description
Dataset Summary
The Defactify_Image_Dataset (A Comprehensive Dataset for Human vs. AI Generated Image Detection) is a high-quality collection of 96,000 images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. Built using the MS… See the full description on the dataset page: https://huggingface.co/datasets/Rajarshi-Roy-research/Defactify_Image_Dataset.PUUM-koa-restoration-camera-trap-dataset
Dataset Card for Koa Associated Biodiversity Camera Trap Dataset
This dataset is aimed at classification of birds visiting planted Acacia koa (koa) trees in the Pu'u Maka'ala Natural Area Reserve (PUUM) on the island of Hawaii (Big Island). The dataset contains full and cropped images collected by camera trap. These images were collected from January 24th to February 25th, 2025.
Dataset Details
This dataset is aimed at classification of birds visiting planted Acacia… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/PUUM-koa-restoration-camera-trap-dataset.NWPU-RESISC45
Dataset Card for "NWPU-RESISC45"
Licensing Information
[CC-BY-SA]
Citation Information
Remote sensing image scene classification: Benchmark and state of the art
@article{cheng2017remote,
title = {Remote sensing image scene classification: Benchmark and state of the art},
author = {Cheng, Gong and Han, Junwei and Lu, Xiaoqiang},
year = 2017,
journal = {Proceedings of the IEEE},
publisher = {IEEE},
volume = 105… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/NWPU-RESISC45.results
TrueVisLies – Results
This dataset contains all raw outputs, extracted fields, semantic similarity scores, and UMAP projections produced in the paper:
True (VIS) Lies: Analyzing How Generative AI Recognizes Intentionality, Rhetoric, and Misleadingness in Visualization Lies
The paper evaluates 16 LLMs, 15 open-weight vision-language models (VLMs), and GPT-5.4 on their ability to (RQ0) detect misleading data visualizations, (RQ1) identify the visualization rhetoric techniques, and… See the full description on the dataset page: https://huggingface.co/datasets/truevislies/results.resisc45
Description
RESISC45 dataset is a publicly available benchmark for Remote Sensing Image Scene Classification (RESISC), created by Northwestern Polytechnical University (NWPU). This dataset contains 31,500 images, covering 45 scene classes with 700 images in each class.
The dataset does not have any default splits. Train, validation, and test splits were based on these definitions here… See the full description on the dataset page: https://huggingface.co/datasets/mteb/resisc45.OpenJev-Vision-Research-v0.1
OpenJev Vision Research v0.1
12,832 image records, with public provenance, original synthetic scenes,
and programmatically derived decision questions.
This is an experimental research dataset for visual posterior learning and
compositional decisions, released with OpenJev.
It is not a reproduction of TypeSafe's proprietary Jev model or training method.
Three separate configurations
Config
Images
What the labels mean
License
synthetic
8,192
Exact… See the full description on the dataset page: https://huggingface.co/datasets/IamBusy/OpenJev-Vision-Research-v0.1.AI-vs-Deepfake-vs-Real-Resized-Aug
🧠 AI vs Deepfake vs Real — Processed Version
This dataset is the result of preprocessing and augmentation applied to the original datasetprithivMLmods/AI-vs-Deepfake-vs-Real.
📘 Overview
This dataset contains a collection of images categorized into three main classes:
🟩 AI-generated
🟥 Deepfake
🟦 Real (authentic human faces)
It is designed for image classification tasks that aim to distinguish between AI-generated, deepfake, and real faces.… See the full description on the dataset page: https://huggingface.co/datasets/chintalaswathi/AI-vs-Deepfake-vs-Real-Resized-Aug.imagenet_1k_resized_256
Dataset Card for "imagenet_1k_resized_256"
Dataset summary
The same ImageNet dataset but all the smaller side resized to 256.
A lot of pretraining workflows contain resizing images to 256 and random cropping to 224x224, this is why 256 is chosen.
The resized dataset can also be downloaded much faster and consume less space than the original one.
See here for detailed readme.
Dataset Structure
Below is the example of one row of data. Note that the… See the full description on the dataset page: https://huggingface.co/datasets/arthtrivedi/imagenet_1k_resized_256.ImageIn_annotations_resized_images
Dataset Card for ImageIn_annotations_resized_images
More Information needed
residuals-fingerprints
RESIDUALS — LiDAR DEM residual fingerprints
39,716 residual images extracted by applying 593 distinct decomposition configurations × 25 upsampling methods to a single Fairfield County, Ohio LiDAR-derived Digital Elevation Model (1500×375 at 3.33 ft/px). Each row pairs a 256×256 PNG of the residual (rendered with the standard RdBu_r colormap, 99th-percentile symmetric clipping) with the algorithm and parameters that produced it, plus a 40-dim signature vector and pre-computed 2D/3D… See the full description on the dataset page: https://huggingface.co/datasets/bshepp/residuals-fingerprints.imagenette-320px-resplit
Imagenette 320px with Fixed Validation and Test Splits
Dataset Description
This dataset is a reproducible, Parquet-based version of the 320px configuration of frgfm/imagenette. Imagenette is a subset of ten readily classified ImageNet classes created for fast experimentation with image-classification methods.
This version preserves the source images, numeric labels, and label metadata. Its only data change is a fixed, stratified division of the original validation… See the full description on the dataset page: https://huggingface.co/datasets/leandrodevai/imagenette-320px-resplit.imagewoof-320px-resplit
ImageWoof 320px with Fixed Validation and Test Splits
Dataset Description
This dataset is a reproducible, Parquet-based version of the 320px configuration of frgfm/imagewoof. ImageWoof is a subset of ten dog-breed classes from ImageNet designed to be more difficult than broad-category image-classification benchmarks.
This version is intended for image classification and confidence-calibration experiments. It introduces two changes to the source dataset:
It… See the full description on the dataset page: https://huggingface.co/datasets/leandrodevai/imagewoof-320px-resplit.eurosat
EuroSAT Image Classification Dataset
This dataset contains the EuroSAT satellite image classification data in parquet format for easy loading and processing.
Dataset Information
Task: Image Classification
Source: EuroSAT Dataset
Classes: 10 land use/land cover classes
Image Size: 64x64 pixels (RGB)
Format: Parquet with embedded images
Splits: train, test
Classes
The dataset contains 10 land use and land cover classes:
ID
Class Name
Description
0… See the full description on the dataset page: https://huggingface.co/datasets/resaro/eurosat.High_Res-vs-Low_Res
High_Res-vs-Low_Res
High_Res-vs-Low_Res is a dataset designed for image classification, distinguishing between high-quality and low-quality images. This dataset includes a diverse collection of 5,016 high-resolution and low-resolution images to enhance classification accuracy and improve the model’s overall efficiency. By providing a well-balanced dataset, it aims to support the development of robust image quality assessment models.
Label Mappings
Mapping of IDs to… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/High_Res-vs-Low_Res.ICDAR2019_cTDaR_TRACKB_resized
Dataset Card for ICDAR2019-cTDaR-TRACKB
This dataset is a resized version of the original cndplab-founder/ICDAR2019_cTDaR, merged with with its supplement cndplab-founder/ICDAR2019_cTDaR_dataset_supplement.
You can easily and quickly load it:
dataset = load_dataset("dvgodoy/ICDAR2019_cTDaR_TRACKB_resized")
DatasetDict({
train: Dataset({
features: ['image', 'width', 'height', 'category', 'label', 'bboxes_table', 'bboxes_cell'],
num_rows: 1200
})
test:… See the full description on the dataset page: https://huggingface.co/datasets/dvgodoy/ICDAR2019_cTDaR_TRACKB_resized.NWPU-RESISC45
Dataset Card for "NWPU-RESISC45"
Licensing Information
[CC-BY-SA]
Citation Information
Remote sensing image scene classification: Benchmark and state of the art
@article{cheng2017remote,
title = {Remote sensing image scene classification: Benchmark and state of the art},
author = {Cheng, Gong and Han, Junwei and Lu, Xiaoqiang},
year = 2017,
journal = {Proceedings of the IEEE},
publisher = {IEEE},
volume… See the full description on the dataset page: https://huggingface.co/datasets/Ling200424/NWPU-RESISC45.radgenome-ct-reshaped-tiny
RadGenome ChestCT Reshaped Tiny Dataset
This dataset contains resized chest CT scans from the RadGenome-ChestCT dataset.
Dataset Details
Original Resolution: 900x900xN
Resized Resolution: 300x300xN
Format: NIfTI (.nii.gz)
Number of Volumes: 253
Space Reduction: ~89% (resized to 1/9th of original spatial dimensions)
Dataset Structure
Each entry contains:
volumename: Name of the CT volume file (string)
anatomy: Anatomical region information (string)
sentence:… See the full description on the dataset page: https://huggingface.co/datasets/nahidhasan/radgenome-ct-reshaped-tiny.ICDAR2019_cTDaR_TRACKA_resized
Dataset Card for ICDAR2019-cTDaR-TRACKA
This dataset is a resized version of the original cndplab-founder/ICDAR2019_cTDaR.
You can easily and quickly load it:
dataset = load_dataset("dvgodoy/ICDAR2019_cTDaR_TRACKA_resized")
DatasetDict({
train: Dataset({
features: ['image', 'width', 'height', 'category', 'label', 'bboxes'],
num_rows: 1200
})
test: Dataset({
features: ['image', 'width', 'height', 'category', 'label', 'bboxes'],
num_rows:… See the full description on the dataset page: https://huggingface.co/datasets/dvgodoy/ICDAR2019_cTDaR_TRACKA_resized.nwpu-resisc45-v1
NWPU-RESISC45
Train/val split manifests extracted from GeoChat_Instruct
for fine-tuning vision-language models on satellite scene classification.
This Hub dataset includes a real image column so the dataset viewer shows the image,
question, and answer aligned row-by-row.
Classes
45 scene categories (e.g. airplane, baseball diamond, beach, bridge, …)
Splits
Split
Samples
train
~25 200
val
~4 725
Format
Columns:
image (previewable… See the full description on the dataset page: https://huggingface.co/datasets/tarek199147/nwpu-resisc45-v1.AI-vs-Deepfake-vs-Real-Resized-Aug
🧠 AI vs Deepfake vs Real — Processed Version
This dataset is the result of preprocessing and augmentation applied to the original datasetprithivMLmods/AI-vs-Deepfake-vs-Real.
📘 Overview
This dataset contains a collection of images categorized into three main classes:
🟩 AI-generated
🟥 Deepfake
🟦 Real (authentic human faces)
It is designed for image classification tasks that aim to distinguish between AI-generated, deepfake, and real faces.
⚙️… See the full description on the dataset page: https://huggingface.co/datasets/riandika/AI-vs-Deepfake-vs-Real-Resized-Aug.NWPU-RESISC45
Dataset Card for "NWPU-RESISC45"
Licensing Information
[CC-BY-SA]
Citation Information
Remote sensing image scene classification: Benchmark and state of the art
@article{cheng2017remote,
title = {Remote sensing image scene classification: Benchmark and state of the art},
author = {Cheng, Gong and Han, Junwei and Lu, Xiaoqiang},
year = 2017,
journal = {Proceedings of the IEEE},
publisher = {IEEE},
volume = 105… See the full description on the dataset page: https://huggingface.co/datasets/yangjiahao-x/NWPU-RESISC45.
