datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
french-lot-department-captioned-photos
Lot Department, France Image Dataset
A collection of high-resolution scenic photographs from the Lot region of France with AI-generated descriptive captions.
Dataset Summary
This dataset contains scenic photographs from three notable locations in France's Lot department: Rocamadour, Autoire, and Padirac. All images were captured using a Sony A6600 camera and are paired with detailed English captions generated by Mistral AI's Pixtral-Large model.
Key Features:… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/french-lot-department-captioned-photos.ne_plant_photos
NE Plant Photos
Plant photographs from the iNaturalist open data
archive, filtered to a New England bounding box (latitude 41 to 48, longitude
-74 to -67), each labelled with the species it records and where it was
observed.
This is the nature portion of
blambert/ne_plant_classes -
photos a vision language model judged to show a plant in the field, rather than
a person, a microscope image, or something manmade - so photos of people and
indoor shots are largely gone, but the… See the full description on the dataset page: https://huggingface.co/datasets/blambert/ne_plant_photos.albi-captioned-photos
Albi, France Image Dataset
A collection of high-resolution scenic photographs from Albi, France with AI-generated descriptive captions.
Dataset Summary
This dataset contains scenic photographs from Albi, France, including the city center, the Toulouse Lautrec museum, and the Sainte-Cécile Cathedral. All images were captured using a Sony A6600 camera and are paired with detailed English captions generated by Mistral AI's Pixtral-Large model.
Key Features:
High-resolution… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/albi-captioned-photos.photo-test
Anime vs. Live Action Film Image Classification
OmarK211/photo-test
A binary image classification dataset designed to distinguish between Anime (Label 0) and Live Action (Label 1) film frames/imagery. Images are prepared as standardized square RGB files with multi-pass synthetic training variants.
Source and task
The images were manually imported from the web into my local computer. Then manually uploaded to Google Colab
Preparation source: 24-679 Image Data… See the full description on the dataset page: https://huggingface.co/datasets/OmarK211/photo-test.QIT-CEMC-FLUTE-PHOTOS
Roles
Roles: canon repo — annot is the source label, kept machine-parseable as the gold for verification and reward parsing; there is no filled reasoning column and this repo is not itself a training view. Derived repos each state their own regime on their own card.
QIT-CEMC 四刃铣刀刃口显微照片(三分类,带 VBmax)
544 张铣刀刃口显微照片,三分类标签(正常 / 磨损 / 毛边卷刃),每张附带上游用显微镜
测量软件量到的 VBmax(后刀面磨损带宽度,毫米)。
照片来自 QIT-CEMC 数据集(Qilu Institute of Technology, Coated End Milling Cutter,… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/QIT-CEMC-FLUTE-PHOTOS.vintage-photography-450k-high-quality-captionsThis is a 450k image datastet focused on photography from the 20th century, and their analog aspect. Many of the images are in high resolution. This dataset currently has 20k images captioned with InternVL2 26B, and is a work in progress (I plan to caption the entire dataset and also have short captions for all of the images, compute is an issue for now).
samuel-and-audrey-photography-metadata-archive
Samuel & Audrey Photography Metadata Archive
This dataset contains a structured metadata archive for the Samuel & Audrey Media Network travel photography collection hosted on SmugMug.
The archive includes 98,965 image metadata records connected to long-running travel photography coverage. Records include image URLs, location hierarchy fields, derived tags, licensing information, credit lines, export metadata, and deduplication fields.
This dataset provides metadata and source URLs… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/samuel-and-audrey-photography-metadata-archive.Photo_Dataset
24-679 (Fall 2026): Stairs and Non-Stair Images
ArinRoths/Photo_Dataset
Photos of stairs and non-stairs scenes, prepared as square RGB images with multiple separately
generated training variants. The goal of this dataset is to classify whether or not stairs are
present in an image.
Source and task
The original dataset contains 32 images that I collected and organized into stairs and
non_stairs folders. There are 16 original stairs images and 16 original non-stairs… See the full description on the dataset page: https://huggingface.co/datasets/ArinRoths/Photo_Dataset.MV-VDB-photos-small
MV-VDB-photos-small
Media Vault - Vector Database Photos (Small)
A curated collection of 11,000 images from various computer vision datasets, designed for testing internal mechanisms in the Media Vault Vector Database system. This is the first small-scale dataset (targeting 10K samples, with NSFW split totaling 11K) for validation and testing purposes.
Dataset Structure
The dataset contains two splits:
sfw: All non-NSFW images (~10,000 images)
x_nsfw: Only NSFW images… See the full description on the dataset page: https://huggingface.co/datasets/SamoXXX/MV-VDB-photos-small.micah-boswell-photography
Micah Boswell photography on Unsplash
Metadata for the most viewed photographs by Micah Boswell (@micahboswell on Unsplash): Dallas at night, neon, wet streets, doors in Peru and Cuba. As of 2026-09-21 the full portfolio is 104 photographs with 53,380,492 views and 362,001 downloads, which places him in the top 10 percent of Unsplash contributors.
photos.jsonl holds one record per photograph: title, Unsplash page, hotlinkable image URL (Unsplash CDN), dimensions, date, likes… See the full description on the dataset page: https://huggingface.co/datasets/socraticstatic/micah-boswell-photography.vintage-photography-450k-high-quality-captionsThis is a 450k image datastet focused on photography from the 20th century, and their analog aspect. Many of the images are in high resolution. This dataset currently has 20k images captioned with InternVL2 26B, and is a work in progress (I plan to caption the entire dataset and also have short captions for all of the images, compute is an issue for now).
