datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
documentation-images
This dataset contains images used in the documentation of HuggingFace's libraries.
HF Team: Please make sure you optimize the assets before uploading them.
My favorite tool for this is https://tinypng.com/.
course-imagesdocumentation-imagesDEH-image-scan-dataimage_dummy\imagesimagenet-1k
Dataset Card for ImageNet
Dataset Summary
ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are… See the full description on the dataset page: https://huggingface.co/datasets/ILSVRC/imagenet-1k.course-imagesgpt-image-2-prompts-datasets
🖼️ GPT Image 2 Prompt Dataset
🖼️ The ultimate GPT Image 2 prompt dataset (5GB+). 15,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for OpenAI's GPT Image 2 model and the resulting generated images. The entire dataset exceeds 5GB and contains 15,000+ images, all structured into a comprehensive dataset.
Due to… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/gpt-image-2-prompts-datasets.images
deepinv/images
Sample images used by DeepInverse examples.
Files in this repo may come from different sources under different licenses. See the README.md in each subfolder for source and license.
imagefolder_with_metadatadocumentation-imagesdiffusers-imagesimagesdocumentation-imagesimagesimage-bank-202601
📂 matitie/image-bank-202601
Dataset Description
This dataset is part of the Top10Fans monthly archival system, designed to store and deliver processed visual assets for blog content, historical reference, and automated publishing workflows. Each month, a new dataset is created to isolate and version image collections by time period.
📅 Timeframe
Month: 202601 (January 2026, UTC)
🎯 Purpose
Serve as a durable, versioned storage layer for… See the full description on the dataset page: https://huggingface.co/datasets/matitie/image-bank-202601.random-imagesfixtures_image_utils\\nimagenet1k-256-wds-latentsThe imagenet1k dataset in the webdataset format
Each image was resized so that the max side resolution is 256, making sure to preserve aspect ratio.
Each image was encoded to latents using the sixteen channel https://huggingface.co/ostris/vae-kl-f8-d16
No cropping was used to encode to latents!
The resulting dataset has images in their original aspect ratio, but much smaller, and encodeded with a vae.
ale-images-qcow2
ALE QEMU runner image
agentslastexam/ale-qemu is the container-side runtime used by the ALE qemu
provider. It packages QEMU, KVM integration, NAT networking, noVNC, and process
supervision. The Ubuntu or Windows guest is supplied separately as
/storage/data.qcow2.
Docker is the container runtime. Dockur is the upstream QEMU-in-Docker project
whose startup and networking stack this image inherits. ALE adds a stable
runner contract around that upstream image.
The image is based on… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/ale-images-qcow2.text-to-image-promptsIf you have questions about this dataset , feel free to ask them on the fusion-discord : https://discord.gg/8TVHPf6Edn
This collection contains sets from the fusion-t2i-ai-generator on perchance.
This datset is used in this notebook: https://huggingface.co/datasets/codeShare/text-to-image-prompts/tree/main/Google%20Colab%20Notebooks
To see the full sets, please use the url "https://perchance.org/" + url
, where the urls are listed below:
_generator
gen_e621
fusion-t2i-e621-tags-1… See the full description on the dataset page: https://huggingface.co/datasets/codeShare/text-to-image-prompts.tiny-imagenet
Dataset Card for tiny-imagenet
Dataset Summary
Tiny ImageNet contains 100000 images of 200 classes (500 for each class) downsized to 64×64 colored images. Each class has 500 training images, 50 validation images, and 50 test images.
Languages
The class labels in the dataset are in English.
Dataset Structure
Data Instances
{
'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=64x64 at 0x1A800E8E190,
'label': 15
}… See the full description on the dataset page: https://huggingface.co/datasets/zh-plus/tiny-imagenet.wds_imagenet_sketchomnibioai-sif-images
OmniBioAI SIF Images 🧬
500+ native ARM64 Singularity (SIF) container images for bioinformatics,
built on NVIDIA DGX (aarch64).
Tool Categories
Category
Tools
Genomics & Alignment
BWA, STAR, HISAT2, Minimap2, Bowtie2
Variant Calling
GATK, DeepVariant, Clair3, Mutect2
RNA-seq
Salmon, Kallisto, DESeq2, edgeR
Single Cell
Seurat, Scanpy, Cell Ranger, Harmony
Epigenomics
MACS2, deepTools, Bismark
Metagenomics
Kraken2, MetaPhlAn, QIIME2
Proteomics… See the full description on the dataset page: https://huggingface.co/datasets/omnibioai/omnibioai-sif-images.TreeOfLife-200M
Dataset Card for TreeOfLife-200M
If you are looking for the original release TreeOfLife-200M dataset, as used in training BioCLIP 2 and presented the paper, please see Revision a8f38b4. The dataset, as presented here, was used to train BioCLIP 2.5 Huge; it completes the dataset cleaning process and resolves an issue where Observation.org occurrences were not included in the training data.
With 233 million images representing 933,798 taxa across the tree of life, TreeOfLife-200M… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/TreeOfLife-200M.JoyAI-Image-OpenSpatial
JoyAI-Image-OpenSpatial
Spatial understanding dataset built on OpenSpatial, used in JoyAI-Image.
The full dataset contains about ~6M multi-turn visual-spatial QA samples across 7 open-source datasets and web data. The open-source datasets contain ARKitScenes, ScanNet, ScanNet++, HyperSim, Matterport3D, WildRGB-D, and Ego-Exo4D. Tasks cover a wide range of spatial understanding capabilities including 3D object grounding, depth ordering, spatial relation reasoning, distance… See the full description on the dataset page: https://huggingface.co/datasets/jdopensource/JoyAI-Image-OpenSpatial.fish-vista
Dataset Card for Fish-Visual Trait Analysis (Fish-Vista)
Note that the '</Use this dataset>' option will only load the CSV files. To download the entire dataset, including all processed images and segmentation annotations, refer to Instructions for downloading dataset and images.
See Example Code to Use the Segmentation Dataset
Figure 1. A schematic representation of the different tasks in Fish-Vista Dataset.
Instructions for downloading dataset… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/fish-vista.imagenetimagenet1k-256-wdsThis is imagenet1k in webdataset format. Images are stored as jpg files. Every image has been resized to a maximum side length of 256. That means that if an image in the original dataset was 1000 by 500, the new size will be 256 by 128. Images with a maximum side length of under 256 were not resized.
The total size of all dataset files is 57.8 GB, there are 1,281,167 rows in the training split and 50,000 rows in the validation split.
