datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
product-photography-v1-tiny-prompts-tasks-collage-filteredmidjourney-v6-recap
Midjourney v6 Recaptioned
~1.2M Midjourney v6 images with captions from three VLMs:
llava: Original LLaVA captions from the source dataset
gemini: Gemini Flash 1.5 captions
qwen3: Qwen3 VL 8B captions
Caption coverage
llava: available for all 1,235,432 images (from original dataset)
gemini and qwen3: available for 1,017,105 images (82.3%)
Source
Based on brivangl/midjourney-v6-llava.
xhs-travel-photos
XHS Travel Photos
Travel photography from Xiaohongshu (Little Red Book) across 19 destinations in China and Southeast Asia.
Dataset Summary
Metric
Count
Notes
2,882
Images
28,005
Total size
7.3 GB
Keyword folders
19
Destinations
Folder
Notes
bali
190
cambodia_angkor
200
chengdu
20
chiang_mai
213
chongqing
194
guilin
225
hainan_sanya
212
indonesia
215
laos
209
malaysia
198
myanmar
122
philippines
210… See the full description on the dataset page: https://huggingface.co/datasets/Rabornkraken/xhs-travel-photos.photoIndonesian-running-photos
Dataset Card for Fotoyu Album Archive
This dataset stores photo and video collections archived from Fotoyu albums using the potoyu-tree-downloader application. It is designed to act as a high-speed Cloudflare-backed CDN for serving static media assets, as well as providing a dataset for image/video classification and machine learning model training.
Dataset Details
Dataset Description
The dataset aggregates scraped photo galleries and video albums… See the full description on the dataset page: https://huggingface.co/datasets/TierKun/Indonesian-running-photos.fermi-lat-weekly-photons
Fermi-LAT weekly photons
This dataset contains Fermi Large Area Telescope all-sky weekly photon files
from mission week w009 through w153, frozen on 2026-08-30. Its 145
configurations correspond one-to-one with the weekly p305_v001 FITS files.
Each Parquet row is an EVENTS row, with the 23 FITS-named columns in their
stored order and shape.
Mission weeks run Thursday through Wednesday in UTC. The first configuration
begins with the science-phase interval on 2008-08-04; w153… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/fermi-lat-weekly-photons.PhotoDoodle2d-photonic-topology
README
This dataset includes the results of a symmetry-based analysis of two-dimensional photonic crystals, spanning 11 distinct symmetry settings, two field polarizations, and five dielectric contrasts. For each of these settings, the dataset includes results for 10 000 randomly generated photonic crystal unit cells. These results assume time-reversal symmetry. In addition, results for time-reversal broken settings are included for 4 of the 11 symmetry settings, at a single… See the full description on the dataset page: https://huggingface.co/datasets/cgeorgiaw/2d-photonic-topology.HDR_Photos_VAE_Training_DNGA collection of HDR images in the DNG format for use in training a HDR VAE.
pexels-photos-janpf
Dataset migrated
This location shoud be considered obsolete. Data has been copied to
https://huggingface.co/datasets/opendiffusionai/pexels-photos-janpf
In a month or so, I should remove the data from here, to be nice to huggingface.
The images there have been renamed to match their md5 checksum. However, a translation table is in this repo, if for some reason you need it.
Downloading
If for some reason, downloading from here is needed, you can use
huggingface-cli… See the full description on the dataset page: https://huggingface.co/datasets/ppbrown/pexels-photos-janpf.chicago-grantpark-photogrammetry
Chicago / Grant Park — Aerial Photogrammetry + Terrestrial Laser Dataset
2,751 aerial photos (45.0 GB) plus the terrestrial laser scans (9.3 GB, 43 stations) over
Grant Park and downtown Chicago, June 2020. CC BY 4.0.
⬇ Download
→ huggingface.co/datasets/Matt1up/chicago-grantpark-photogrammetry
Browse the Files tab and take what you want — no account needed. The 54 GB of imagery and
laser scans lives there because GitHub won't host files that size;… See the full description on the dataset page: https://huggingface.co/datasets/Matt1up/chicago-grantpark-photogrammetry.terminusresearch-photo-aesthetics
Photo Aesthetics — Webshart
This is the canonical, maintained location for the former terminusresearch/photo-aesthetics dataset. The migration completed on August 19, 2026, with 30,032 image samples repackaged into 371 indexed Webshart shards. Every sample now includes both the original CogVLM caption and a newer GLM-5.3-Flash caption.
The legacy repository now contains only a relocation notice. Its original tar archives, parquet file, and prior history were removed after this… See the full description on the dataset page: https://huggingface.co/datasets/webshart/terminusresearch-photo-aesthetics.Duskfallcrew_Unsplash_PhotographyThis was direct uploads from the photographer.
We've long since retired from AI, but if you find our content useful even long after please feel free to send us a KoFI
Otherwise feel free to partake in our affiliate links
Membership / Ko-Fi
Runpod
VastAI
Photographer
cmp-vialitix-photos-2025
Panoramax sig14 photos
Attribution — Ce jeu de données contient des images issues de la plateforme
Panoramax de l'IGN (Institut national de l'information
géographique et forestière). Les images d'origine sont publiées sous
Licence Ouverte / Open License 2.0 (Etalab)
par leur producteur sig14 (Service d'Information Géographique du Calvados) et l'IGN.
Street-view photos from the IGN Panoramax instance, scraped from the user sig14
(Service d'Information Géographique du Calvados… See the full description on the dataset page: https://huggingface.co/datasets/calvadosdep/cmp-vialitix-photos-2025.cosb-photon-events
COS-B photon events
python -m venv .venv && .venv/bin/pip install datasets huggingface_hub pyarrow
This dataset contains the archive-natural event products for all 65 COS-B
pointings served by HEASARC. Each configuration is one source event file;
pointing 37, targeted on M31, is used in the examples below.
Cite H. A. Meyer-Hasselwander et al. (1986), Explanatory Supplement to the
COS-B Final Database, Proc. Cosmic Ray Conf., La Jolla, ESA Publication.
No explicit dataset… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/cosb-photon-events.pexels-photos-janpf
Images:
There are approximately 130K images, borrowed from pexels.com.
Thanks to those folks for curating a wonderful resource.
There are millions more images on pexels. These particular ones were selected by
the list of urls at https://github.com/janpf/self-supervised-multi-task-aesthetic-pretraining/blob/main/dataset/urls.txt .
The filenames are based on the md5 hash of each image.
Download From here or from pexels.com: You choose
For those people who like… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/pexels-photos-janpf.vFab-1.0-Photolithography-ML-surrogate-Tensor-Dataset
vFab-1.0-Photolithography-ML-surrogate-Tensor-Dataset
This dataset contains generated physics simulations for semiconductor photolithography, formatted as Native NCHW .npy arrays for direct ingestion into PyTorch models.
Dataset Structure
The data is split into two simulation phases:
1. Aerial Image Dataset (aerial_image_dataset/)
Inputs (X_input/): Shape [4, 512, 512]. Contains:
Channel 0: Binary Layout Mask (Vertical/Horizontal Gratings, Contact Arrays, etc.)… See the full description on the dataset page: https://huggingface.co/datasets/nk-21-mit/vFab-1.0-Photolithography-ML-surrogate-Tensor-Dataset.img-256-photo-2
Dataset Card for "img-256-photo-2"
More Information needed
improved-flux-prompts-photoreal-portrait
Photo Portrait Prompt Dataset for FLUX
Overview
This dataset contains a curated collection of prompts specifically designed for generating photo portraits using FLUX.1, an advanced text-to-image model. These prompts are crafted to produce high-quality, lifelike portraits by leveraging sophisticated prompting techniques and best practices.
Latest Version
Improved on October 3, 2024.
This version has undergone curation and improvement. What is new?
Cleaned up… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/improved-flux-prompts-photoreal-portrait.tree-minnetonka-photogrammetry
Single Tree — High-Density Photogrammetry Dataset
812 photos of one mature deciduous tree, flown from 0.5 to 7.9 m above the ground. 15.1 GB.
807 align, and the solved camera poses ship with it in COLMAP format — point a Gaussian
splatting pipeline straight at it, no structure-from-motion run required. CC BY 4.0.
⬇ Download
→ huggingface.co/datasets/Matt1up/tree-minnetonka-photogrammetry
Browse the Files tab and take what you want — no account… See the full description on the dataset page: https://huggingface.co/datasets/Matt1up/tree-minnetonka-photogrammetry.family-photo-thumbnailsshelf-photos-batch1
shelf-photos-batch1
Working dataset for the shelf-monitoring pipeline: bootstrap labeling and
fine-tuning data management for Geraldine/rf-detr-nano-bookshelf.
Layout
photos/ — 25 of the library's own shelf photos, untouched (no crops, no upscaling)
original_images_library/ — 285 high-res images from
llabres/library-dataset (MIT),
used for continuous fine-tuning (domain shift: real library stacks)
dataset/Bookshelf-recognition-2.v1i.coco.zip — COCO export of… See the full description on the dataset page: https://huggingface.co/datasets/Geraldine/shelf-photos-batch1.japanese-photos
Japan Diverse Images Dataset
Overview
This dataset is a comprehensive collection of high-quality images capturing the diverse aspects of Japan, including urban landscapes, natural scenery, historical sites, contemporary art, everyday life, and culinary experiences. It is designed to provide a rich and varied representation of Japan for AI training purposes.
Note that the photos were taken by myself in the 2020s, mainly from 2022 to 2024, with some exceptions.… See the full description on the dataset page: https://huggingface.co/datasets/ThePioneer/japanese-photos.my-photosphoton47-trailers
Photon 47 bilingual trailers
This public media dataset contains the matched English and Chinese Photon 47
gameplay trailers used for promotion and review. The current editions follow
the same 289-second, 8,670-frame edit across all eight game modes; only
localized callout and subtitle text differs.
The release inventory is intentionally small and exact:
photon47-trailer.en.mp4
photon47-trailer.zh.mp4
photon47-trailer.en.vtt
photon47-trailer.zh.vtt
trailer-manifest.json… See the full description on the dataset page: https://huggingface.co/datasets/ryan-superman/photon47-trailers.getting-started-labeled-photos
Dataset Card for predicted_labels
These photos are used in the FiftyOne getting started webinar. The images have a prediction label where were generated by
self-supervised classification through a OpenClip Model.
https://github.com/thesteve0/fiftyone-getting-started/blob/main/5_generating_labels.py
They were then manually cleaned to produce the ground truth label.
https://github.com/thesteve0/fiftyone-getting-started/blob/main/6_clean_labels.md
They are 300 public domain photos… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/getting-started-labeled-photos.photonics_2d_120_120_v0This dataset is part of the EngiBench toolkit: https://github.com/IDEALLab/EngiBench.
photonics_2d_120_120_v1
Photonics 2D (120x120) - Wavelength Demultiplexer (v1)
Inverse-design dataset for the EngiBench photonics2d problem: 2D photonic wavelength
demultiplexers optimized with a finite-difference frequency-domain (FDFD) solver (ceviche).
Each design routes light at two wavelengths to two separate output ports.
v1
v1 makes simulate and optimize mutually consistent: designs are stored as physical (projected)
densities in [0, 1], and simulate evaluates them as-is, so the… See the full description on the dataset page: https://huggingface.co/datasets/IDEALLab/photonics_2d_120_120_v1.barney-photography-deliverablespublic_flickr_photos_license_1
119893266 photos from flickr (https://www.flickr.com/creativecommons/by-nc-sa-2.0/)
all photos are under license id = 1 name=Attribution-NonCommercial-ShareAlike License url=https://creativecommons.org/licenses/by-nc-sa/2.0/
