datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
efficientnet-v2-l-adv-dataset
Perturb Adversarial Images
Verified adversarial examples for efficientnet_v2_l (torchvision/EfficientNet_V2_L_Weights.IMAGENET1K_V1), produced by the
Perturb network. Each row is one clean image together with all of its
verified adversarial versions: images that are imperceptibly different from the original
(L∞ ≤ 0.03 in [0,1] pixel scale) yet change the model's top-1 prediction.
This dataset grows continuously. New rows are appended as the network produces them and uploaded in… See the full description on the dataset page: https://huggingface.co/datasets/perturb-ai/efficientnet-v2-l-adv-dataset.LADI-v2-dataset
Dataset Card for LADI-v2-dataset
Dataset Summary : v2
The LADI-v2 dataset is a set of aerial disaster images captured and labeled by the Civil Air Patrol (CAP). The images are geotagged (in their EXIF metadata). Each image has been labeled in triplicate by CAP volunteers trained in the FEMA damage assessment process for multi-label classification; where volunteers disagreed about the presence of a class, a majority vote was taken. The classes are:
bridges_any… See the full description on the dataset page: https://huggingface.co/datasets/MITLL/LADI-v2-dataset.ScreenSpot-v2
Dataset Card for ScreenSpot-V2
This is a FiftyOne dataset with 1272 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/ScreenSpot-v2")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/ScreenSpot-v2.autotrain-data-logo-identifier-v2-short
AutoTrain Dataset for project: logo-identifier-v2-short
Dataset Description
This dataset has been automatically processed by AutoTrain for project logo-identifier-v2-short.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<100x100 RGB PIL image>",
"target": 98
},
{
"image": "<100x100 RGB PIL image>",
"target": 3… See the full description on the dataset page: https://huggingface.co/datasets/fsuarez/autotrain-data-logo-identifier-v2-short.toothbrush-v2-dataset
Toothbrushing Detection Dataset (v2)
Video and image data for detecting toothbrushing behavior, collected for a Raspberry Pi Zero 2W toothbrush-detection project (toothbrush_v2). A single-class object detector is trained on this data to output [x, y, w, h, confidence] for the toothbrush in frame.
Dataset structure
Files are packed into tar shards (rather than uploaded individually) to stay within the Hub's per-repo file-count guidelines. To reconstruct the… See the full description on the dataset page: https://huggingface.co/datasets/Schrodin-purrrrr/toothbrush-v2-dataset.underworld_dataset_v2
_ _ _ _______ ___________ _ _ ___________ _ ______
| | | | \ | | _ \ ___| ___ \ | | || _ | ___ \ | | _ \
| | | | \| | | | | |__ | |_/ / | | || | | | |_/ / | | | | |
| | | | . ` | | | | __|| /| |/\| || | | | /| | | | | |
| |_| | |\ | |/ /| |___| |\ \\ /\ /\ \_/ / |\ \| |___| |/ /
\___/\_| \_/___/ \____/\_| \_|\/ \/ \___/\_| \_\_____/___/
Underworld Dataset v2
Generated with WEBXOS UNDERWORLD LANDSCAPE GENERATOR… See the full description on the dataset page: https://huggingface.co/datasets/webxos/underworld_dataset_v2.japanese-aerial-fireworks-v2
🎆 NEW: Curated 1,000 Wide Pack (Commercial License)
For commercial AI/ML training, check out the Hanabi AI Dataset v1: Wide Pack — Curated 1,000 — a carefully selected subset with detailed structured annotations:
✅ 1,000 hand-curated 4K images (vs 2,557 raw images here)
✅ Structured AI annotations (composition, mood, color, EXIF, English notes)
✅ Sample PyTorch loader, attribute filter, caption generator
✅ Perpetual Commercial License (Japanese law)
✅ Optimized for Stable… See the full description on the dataset page: https://huggingface.co/datasets/dfhjs2577/japanese-aerial-fireworks-v2.Recraft-V2_t2i_human_preference
Rapidata Recraft-V2 Preference
This T2I dataset contains over 195k human responses from over 47k individual annotators, collected in just ~1 Day using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating Recraft-V2 across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Recraft-V2_t2i_human_preference.Ideogram-V2_t2i_human_preference
Rapidata Ideogram-V2 Preference
This T2I dataset contains over 195k human responses from over 42k individual annotators, collected in just ~1 Day using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating Ideogram-V2 across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Ideogram-V2_t2i_human_preference.food-category-classification-v2.0
Dataset for project: food-category-classification-v2.0
Dataset Description
This dataset for project food-category-classification-v2.0 was scraped with the help of a bulk google image downloader.
Dataset Structure
Dataset Fields
The dataset has the following fields (also called "features"):
{
"image": "Image(decode=True, id=None)",
"target": "ClassLabel(names=['Bread', 'Dairy', 'Dessert', 'Egg', 'Fried Food', 'Fruit', 'Meat', 'Noodles', 'Rice'… See the full description on the dataset page: https://huggingface.co/datasets/Kaludi/food-category-classification-v2.0.SynthCheX-75K-v2
SynthCheX-75K
SynthCheX-75K is released as a part of the CheXGenBench paper. It is a synthetic dataset generated using Sana (0.6B) [1] fine-tuned on chest radiographs. Sana (0.6B) establishes the SoTA performance on the CheXGenBench benchmark.
The dataset contains 75,649 high-quality image-text samples along with the pathological annotations.
Filtration Process for SynthCheX-75K
Generative models can lead to both high and low-fidelity generations on different subsets… See the full description on the dataset page: https://huggingface.co/datasets/raman07/SynthCheX-75K-v2.safemaize-v2
SafeMaize v2
We use this preliminary public-source dataset for maize screening experiments.
The export contains 46,143 distinct images, including
40,385 core classification images. Files retain their original
bytes and recorded frame-selection rules.
Core class
Images
nlb_tlb_like
20,661
healthy
14,108
faw_feeding_injury
5,616
Core screening task split
Images
train
28,269
val
4,039
calibration
4,038
test
4,039
Core primary source… See the full description on the dataset page: https://huggingface.co/datasets/sathiiii/safemaize-v2.Deepfake-vs-Real-v2
Deepfake-vs-Real-v2
Deepfake-vs-Real-v2 is a dataset designed for image classification, distinguishing between deepfake and real images. This dataset includes a diverse collection of high-quality deepfake images to enhance classification accuracy and improve the model’s overall efficiency. By providing a well-balanced dataset, it aims to support the development of more robust deepfake detection models.
Label Mappings
Mapping of IDs to Labels: {0: 'Deepfake', 1:… See the full description on the dataset page: https://huggingface.co/datasets/dappai/Deepfake-vs-Real-v2.M-Attack-V2-Adversarial-Samples
M-Attack-V2 Adversarial Samples
Adversarial image samples generated by M-Attack-V2, from the paper:
Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting
arXiv:2602.17645 | Project Page | Code
Dataset Structure
├── epsilon_8/ # 100 adversarial images (ε = 8/255)
│ ├── 0.png
│ ├── 1.png
│ ├── ...
│ └── metadata.csv
└── epsilon_16/ # 100 adversarial images (ε = 16/255)
├── 0.png
├── 1.png
├── ...
└──… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-LLM/M-Attack-V2-Adversarial-Samples.Dog_Emotion_Dataset_v2
Dataset Card for "Dog_Emotion_Dataset_v2"
The Dataset is based on a kaggle dataset
Label and its Meaning
0 : sad"
1 : angry"
2 : relaxed"
3 : happy"
astrobridge-yse-test-dataset-v2
AstroBridge YSE external test dataset v2
This dataset contains 266 object-disjoint, spectroscopically labeled YSE DR1 transients. The broad-class counts are SN II: 71, SN Ia: 180, SN Ibc: 15.
V2 shortens the YSE forced-photometry time coverage to resemble the alert-photometry coverage of the AstroBridge BTS training dataset. For each object, it retains the smallest inclusive time interval that contains every positive measurement with flux/uncertainty at least 5 and at least five… See the full description on the dataset page: https://huggingface.co/datasets/BuildNg/astrobridge-yse-test-dataset-v2.autotrain-data-multifamily_v2
AutoTrain Dataset for project: multifamily_v2
Dataset Description
This dataset has been automatically processed by AutoTrain for project multifamily_v2.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<500x333 RGB PIL image>",
"target": 33
},
{
"image": "<500x667 RGB PIL image>",
"target": 11
}]
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/lineups-io/autotrain-data-multifamily_v2.iNaturalist_v2
Dataset Card for Dataset Name
This dataset is comprised of 1,079 observations that were posted on the iNaturalist app. iNaturalist is a website and mobile app that 'aims to
provide a crowd-sourced identification system' for plants, insects, and animals.
Dataset Details
Dataset Description
For each of the 1,079 observations included in this dataset, there is information about the quality of the associated image (quality_grade), a species label… See the full description on the dataset page: https://huggingface.co/datasets/ba188/iNaturalist_v2.Deepfake-vs-Real-v2
Deepfake-vs-Real-v2
Deepfake-vs-Real-v2 is a dataset designed for image classification, distinguishing between deepfake and real images. This dataset includes a diverse collection of high-quality deepfake images to enhance classification accuracy and improve the model’s overall efficiency. By providing a well-balanced dataset, it aims to support the development of more robust deepfake detection models.
Label Mappings
Mapping of IDs to Labels: {0: 'Deepfake', 1:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Deepfake-vs-Real-v2.ScreenSpot-v2
Dataset Card for ScreenSpot-V2
This is a FiftyOne dataset with 1272 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/ScreenSpot-v2")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/ZhuOnR/ScreenSpot-v2.LSUN_bedroom_VQA_v2
Dataset Card for "CSUN_bedroom_VQA_feliu_v2"
waikato_aerial_2017_synthetic_v2
Waikato Aerial Imagery 2017 Synthetic Data v2
This is a synthetic dataset generated using a sample taken from the original classification dataset residing at https://datasets.cms.waikato.ac.nz/taiao/waikato_aerial_imagery_2017/. You can find additional dataset information using the provided URL. This version (v2) has been generated using slightly altered prompts compared to v1.
Generation Params
Inference Steps: 60Images generated per prompt: 50 (1000 images per… See the full description on the dataset page: https://huggingface.co/datasets/dinushiTJ/waikato_aerial_2017_synthetic_v2.screenspot_v2_w_gui_actor
Dataset Card for Voxel51/ScreenSpot-v2
This is a FiftyOne dataset with 1272 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("harpreetsahota/screenspot_v2_w_gui_actor")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/screenspot_v2_w_gui_actor.deepfake-detection-dataset-v2
Deepfake Detection Dataset V2
This dataset contains images and detailed explanations for training and evaluating deepfake detection models. It includes original images, manipulated images, confidence scores, and comprehensive technical and non-technical explanations.
Dataset Structure
The dataset consists of:
Original images
CAM visualization images
CAM overlay images
Comparison images
Labels (real/fake)
Confidence scores
Image captions
Technical and non-technical… See the full description on the dataset page: https://huggingface.co/datasets/saakshigupta/deepfake-detection-dataset-v2.brain-tumour-v2Nexora-vision-dataset-v2-medium
Nexora Vision Dataset v2 Medium
The Nexora Vision Dataset v2 Medium is a scalable, mixed-resolution image dataset designed for generative AI experimentation, diffusion model workflows, and computer vision research.
Developed and curated by ArkDevLabs / ArkAiLab (ADL).
Official Website: https://arkdevlabs.com
Dataset Summary
Nexora Vision Dataset v2 Medium contains 9,236 curated images packaged in both:
Raw image format
Optimized Parquet format
This release prioritizes:… See the full description on the dataset page: https://huggingface.co/datasets/ArkAiLab-Adl/Nexora-vision-dataset-v2-medium.stickers-binary-v2-cleaned
Stickers Binary v2 — Cleaned
Binary SFW/NSFW sticker classification dataset. This version has been cleaned
of likely label errors using cross-validated out-of-fold model predictions
combined with cleanlab's
find_label_issues.
Structure
This dataset has exactly two columns:
Column
Type
Description
image
image
The sticker image, 256x256, letterboxed (see below).
label
int64
0 = SFW, 1 = NSFW.
Class distribution
Split
Count… See the full description on the dataset page: https://huggingface.co/datasets/Pankaj8922/stickers-binary-v2-cleaned.diffusion_model_assessment_v2
Dataset Card for generated_flowers_with_embeddings
This is a FiftyOne dataset with 286 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("CarloColumbo/diffusion_model_assessment_v2")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/CarloColumbo/diffusion_model_assessment_v2.graph_dataset_generated_v2
Graph Dataset - Image & LabelMe & OBB Annotation (Train/Val Split)
Dataset Overview
Comprehensive graph/chart detection dataset with ground truth LabelMe polygon annotations and OBB (Oriented Bounding Box) data, split into training and validation sets.
Total examples: 35561 image-annotation pairs
Train: 28448 (80.0%)
Validation: 7113 (20.0%)
Total size: 2134.30 MB
Language: Khmer (km)
Document types: Graph/Chart documents
Ground truth: LabelMe polygon annotations… See the full description on the dataset page: https://huggingface.co/datasets/vichetkao/graph_dataset_generated_v2.mnist-cleaned-joscha-idk-label-v2
Dataset Card for 2025.11.23.16.31.34.701243
This is a FiftyOne dataset with 281 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("joscha-s/mnist-cleaned-joscha-idk-label-v2")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/joscha-s/mnist-cleaned-joscha-idk-label-v2.
