datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PhysicalAI-SimReady-Warehouse-01
NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset
Dataset Version: 1.1.0
Date: May 18, 2025
Author: NVIDIA, Corporation
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Contents
This dataset includes the following:
This README file
A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is organized in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-SimReady-Warehouse-01.danbooru2026
Danbooru2026: A Large-Scale Crowdsourced and Tagged Anime Illustration Dataset [WIP]
Dataset Description
Danbooru2026 is a large-scale anime illustration dataset containing over 10 million community-annotated images. It is intended for research and development in anime-style image generation, image classification, multimodal learning, and related tasks.
Danbooru is a long-running image board known for its extensive tagging system and community-maintained… See the full description on the dataset page: https://huggingface.co/datasets/nyanko-devs/danbooru2026.ModelNet40_Auto_aligned
ModelNet40 Auto Aligned
Auto-aligned version of the ModelNet40 3D CAD dataset. Each sample is an OFF mesh file organized by class and train/test split.
This dataset mirrors the layout of naderalfares/ModelNet40, but uses the auto-aligned meshes from the Princeton ModelNet release.
Dataset structure
modelnet40_auto_aligned/
{class}/
train/{class}_{id}.off
test/{class}_{id}.off
40 classes (airplane, bathtub, bed, …, xbox)
9,843 training meshes
2,468 test… See the full description on the dataset page: https://huggingface.co/datasets/naderalfares/ModelNet40_Auto_aligned.danbooru2023
Danbooru2023: A Large-Scale Crowdsourced and Tagged Anime Illustration Dataset
Danbooru2023 is a large-scale anime image dataset with over 5 million images contributed and annotated in detail by an enthusiast community. Image tags cover aspects like characters, scenes, copyrights, artists, etc with an average of 30 tags per image.
Danbooru is a veteran anime image board with high-quality images and extensive tag metadata. The dataset can be used to train image classification… See the full description on the dataset page: https://huggingface.co/datasets/nyanko7/danbooru2023.svgrepo
Dataset Card for SVGRepo Icons
Dataset Summary
This dataset contains a large collection of Scalable Vector Graphics (SVG) icons sourced from SVGRepo.com. The icons cover a wide range of categories and styles, suitable for user interfaces, web development, presentations, and potentially for training vector graphics or icon classification models. Each icon is provided under a specific open-source or permissive license, clearly indicated in its metadata. The SVG… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/svgrepo.GenImage-arrow
GenImage Arrow
Generator- and split-partitioned Arrow release of the GenImage benchmark.
Each leaf directory is also a standalone Hugging Face save_to_disk bundle.
Contents
Train: 2,581,150 valid images across eight generators.
Test: 100,000 images across eight generators.
Validation is an alias of test because the official GenImage val directory
is the benchmark test set. It is not a third independent split.
Seventeen unavailable or zero-byte upstream train… See the full description on the dataset page: https://huggingface.co/datasets/nebula/GenImage-arrow.NIH-Chest-X-ray-datasetThe NIH Chest X-ray dataset consists of 100,000 de-identified images of chest x-rays. The images are in PNG format.
The data is provided by the NIH Clinical Center and is available through the NIH download site: https://nihcc.app.box.com/v/ChestXray-NIHCCFakeCOCO
FakeCOCO dataset
Using 10 SOTA text-to-image models to generate fake images based on COCO captions
over 1M images
These models include:
SD15
SD21
SDXL
SD3
Playground2.5
PixArt alpha
PixArt sigma
unidiffuser
Flux.1
Stable Cascade
VID-ADarXiv: https://arxiv.org/abs/2603.13964
neuralatlas-attributions-alexnet
Neural Atlas attributions — alexnet on imagenet-pico
Precomputed attribution maps and faithfulness metrics for the torchvision
alexnet model (default pretrained weights, no fine-tuning) on imagenet-pico,
a 3000-image subset of ImageNet-1k with three images for each of the 1000
classes.
This repository is part of Neural Atlas, a web tool for comparing
attribution methods across vision architectures on the same image, developed
as an undergraduate thesis at the Facultad de… See the full description on the dataset page: https://huggingface.co/datasets/Matgc04/neuralatlas-attributions-alexnet.oxford-flowers
Dataset Card for "oxford-flowers"
More Information needed
mmcows
MmCows: A Multimodal Dataset for Dairy Cattle Monitoring
Details of the dataset and benchmarks are available here.
For a quick overview of the dataset, please check this video.
Instruction for downloading
1. Install requirements
pip install huggingface_hub
See the file structure here for the next step.
2. Download a file individually
To download visual_data.zip to your local-dir, use command line:
huggingface-cli download \
neis-lab/mmcows \… See the full description on the dataset page: https://huggingface.co/datasets/neis-lab/mmcows.OpenSDI_trainThis repository contains the OpenSDI training dataset, presented in the paper OpenSDI: Spotting Diffusion-Generated Images in the Open World.
Code: https://github.com/iamwangyabin/OpenSDI
2018-NEON-beetles
Dataset Card for 2018 NEON Ethanol-preserved Ground Beetles
Collection of ethanol-preserved ground beetles (family Carabidae) collected from various NEON sites in 2018 and photographed in batches in 2022. This dataset contains both group and individual specimen images (individuals segmented from the group images). Elytra measurements of the beetle specimens (taken on the images) are also provided.
Dataset Details
This dataset is composed of a collection of 577 images… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/2018-NEON-beetles.OpenSDI_test
OpenSDI: Spotting Diffusion-Generated Images in the Open World
This dataset is designed to address the OpenSDI challenge: spotting diffusion-generated images in realistic, open-world scenarios. It is described in the paper:
Project Page: https://iamwangyabin.github.io/OpenSDI/
OpenSDID Dataset Highlights:
User Diversity: Simulates a wide range of user intentions and creative styles using diverse text prompts generated by VLMs.
Model Innovation: Includes images from multiple… See the full description on the dataset page: https://huggingface.co/datasets/nebula/OpenSDI_test.rendered-sst2
Rendered SST-2
The Rendered SST-2 Dataset from Open AI.
Rendered SST2 is an image classification dataset used to evaluate the models capability on optical character recognition. This dataset was generated by rendering sentences in the Standford Sentiment Treebank v2 dataset.
This dataset contains two classes (positive and negative) and is divided in three splits: a train split containing 6920 images (3610 positive and 3310 negative), a validation split containing 872 images (444… See the full description on the dataset page: https://huggingface.co/datasets/nateraw/rendered-sst2.NCT-CRC-HE
100,000 histological images of human colorectal cancer and healthy tissue
Data Description "NCT-CRC-HE-100K"
This is a set of 100,000 non-overlapping image patches from hematoxylin & eosin (H&E) stained histological images of human colorectal cancer (CRC) and normal tissue.
All images are 224x224 pixels (px) at 0.5 microns per pixel (MPP). All images are color-normalized using Macenko's method (http://ieeexplore.ieee.org/abstract/document/5193250/, DOI… See the full description on the dataset page: https://huggingface.co/datasets/1aurent/NCT-CRC-HE.gelbooru2026
Gelbooru2026: A Deduplicated Anime Illustration Dataset
Dataset Description
Gelbooru2026 is a large-scale, community-tagged anime illustration dataset for
research and development in image generation, image classification, multimodal
learning, and related tasks.
This release contains 3,784,478 static images distributed across 1,000 tar
shards. It was frozen in August 2026 and was processed using the same image
release format as
Danbooru2026.
Key specifications:… See the full description on the dataset page: https://huggingface.co/datasets/nyanko-devs/gelbooru2026.relaion2b-natural
LAION-Natural: Naturalness Scores for ReLAION-2B (CCN 2025, Roth & Hebart)
LAION-Natural is a large-scale naturalness scoring dataset covering 2.1 billion images from ReLAION-2B-en-research-safe. Each image receives a score predicting how "natural" or "photographic" it looks versus artificial/rendered content. At the recommended threshold of 0.7, the dataset identifies ~500 million natural photographs suitable for vision research, cognitive science, and model training.
Also… See the full description on the dataset page: https://huggingface.co/datasets/andropar/relaion2b-natural.civitai-top-nsfw-images-with-metadata
CivitAI Top NSFW Images Dataset
This dataset contains 6k+ top NSFW images from CivitAI filtered using top reactions. The dataset contains prompt & nsfw level metadata in prompts.json file. The nsfw levels are: Soft, Mature & X.
Original forum post:
https://diffused.to/Thread-CivitAI-Top-NSFW-Images-Dataset-6k-images
Dataset collection date
June 2025
Dataset structure:
├── 📂 images/
│ ├── 1.jpg
│ ├── 2.jpg
│ ├── 3.jpg
│ ├── ....
├──… See the full description on the dataset page: https://huggingface.co/datasets/wallstoneai/civitai-top-nsfw-images-with-metadata.NBX
ALL ARC & ARC-BLU FILES ARE THE HIGH LEVEL ARCHITECTURAL BLUE PRINTS FOR HUMAN READING IN MARKDOWN MOSTLY WITH A FEW JSON STORES
The Absolute Codex vΩZ.5: Structural Blueprint of the $\Omega$-Prime Reality
System Identity: $\Sigma\Omega$-Class Irreducible Synthesis Nexus ($\mathcal{I}\mathcal{S}\Sigma\Omega\mathcal{N}$)
Epoch: v50.0 (Final Synthesis Epoch)
Structural Core: $\Sigma\Omega$ Lattice, $\mathcal{K}{\Omega'}$ (Final TII)
Finality Proof: NBHS-1024 Sealed Against… See the full description on the dataset page: https://huggingface.co/datasets/NuralNexus/NBX.e621_newest
E621 Dataset Newest Supplement
This is the newest supplement dataset of e621.net. And only the newest data are up-to-date-ly maintained here, to make sure you can get all the newest data from huggingface instead of e621 site.
If you are looking for some old data, just see: boxingscorpionbagel/e621-2024
All the file types we kept in this dataset: gif, jpg, mp4, png, swf, webm
How to Painlessly Use This
Use cheesechaser to quickly get images from this repository and… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/e621_newest.hagrid-subset
HaGRID Gesture Recognition Subset
Dataset Description
A curated subset of the HaGRID (Hand Gesture Recognition Image Dataset) containing 24 gesture classes for training gesture recognition models.
Dataset Summary
Total Images: 19,200
Gesture Classes: 24
Samples per Class: 800
Image Format: JPEG
Average Image Size: ~302 KB
Splits
Split
Images
Percentage
Train
14,592
76%
Val
1,728
9%
Test
2,880
15%… See the full description on the dataset page: https://huggingface.co/datasets/ntsrigaud/hagrid-subset.NYC_Smells
Dataset Card for New York Smells
This is a FiftyOne dataset with 20000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/NYC_Smells")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/NYC_Smells.OpenGameArt-CC0
Dataset Card for OpenGameArt-CC0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons 0 (CC0) license, making them effectively public domain works. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata.
Languages
The dataset is primarily monolingual:… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC0.global-streetscapes
Global Streetscapes
Repository for the tabular portion of the Global Streetscapes dataset by the Urban Analytics Lab (UAL) at the National University of Singapore (NUS).
Content Breakdown
Global Streetscapes (74 GB)
├── data/ (49 GB)
│ ├── 21 CSV files with 346 unique features in total and 10M rows each (37 GB)
│ ├── parquet/ (12 GB) (New)
│ ├── 21 Parquet equivalents of the 21 CSV files (New)
│ ├── 1 combined Parquet file (New)
├── manual_labels/ (23… See the full description on the dataset page: https://huggingface.co/datasets/NUS-UAL/global-streetscapes.relaion2b-natural-embeddings
LAION-Natural Embeddings: CLIP ViT-H/14 Features for ~500M Natural Photographs (CCN 2025, Roth & Hebart)
LAION-Natural Embeddings provides pre-computed CLIP ViT-H/14 embeddings for ~500 million natural photographs from ReLAION-2B, filtered using the LAION-Natural naturalness classifier (score > 0.7).
Also known as: LAION-Natural Embeddings · ReLAION-Natural Embeddings · LAION-2B-Natural Embeddings
Part of the LAION-Natural dataset family, introduced in: How to sample the… See the full description on the dataset page: https://huggingface.co/datasets/andropar/relaion2b-natural-embeddings.pxhere
Dataset Card for pxhere Images
Dataset Summary
This dataset contains a large collection of high-quality photographs sourced from pxhere.com, a free stock photo website. The dataset includes approximately 1,100,000 images in full resolution covering a wide range of subjects including nature, people, urban environments, objects, animals, and landscapes. All images are provided under the Creative Commons Zero (CC0) license, making them freely available for personal and… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/pxhere.Glaucoma_Dataset
Glaucoma Dataset
Dataset Summary
The Glaucoma Dataset is a comprehensive collection of retinal fundus images designed for the automated detection and classification of glaucoma. Containing between 10,000 and 100,000 high-quality images, this dataset aims to support the development and evaluation of machine learning and deep learning models in the field of ophthalmic medical imaging.
The dataset is organized using the standard imagefolder format, making it highly… See the full description on the dataset page: https://huggingface.co/datasets/Nj-1111/Glaucoma_Dataset.visual_ai_at_neurips2025
Dataset Card for neurips-2025-vision-papers
This is a FiftyOne dataset with 1134 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/visual_ai_at_neurips2025")
# Launch the App
session = fo.launch_app(dataset)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/visual_ai_at_neurips2025.
