datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Defactify_Image_Dataset
Defactify_Image_Dataset
This dataset is associated with the paper A Comprehensive Dataset for Human vs. AI Generated Image Detection.
📝 Dataset Description
Dataset Summary
The Defactify_Image_Dataset (A Comprehensive Dataset for Human vs. AI Generated Image Detection) is a high-quality collection of 96,000 images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. Built using the MS… See the full description on the dataset page: https://huggingface.co/datasets/Rajarshi-Roy-research/Defactify_Image_Dataset.text-2-image-human-preferences-2m
Text-to-image human preferences: 2M votes across 30 models
This dataset contains the complete voting record behind the
Datapoint Image Bench
leaderboard: 2,161,160 validated pairwise votes — exactly 10 for each of
216,116 image pairs. The votes compare 30 text-to-image models in a complete
round-robin on 500 prompts, judged by annotators from over 200 countries.
Every vote includes the annotator's trust score at the time the vote was
cast.
Built on the Datapoint annotation… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-human-preferences-2m.seatizen_atlas_image_dataset
Seatizen Atlas Image Dataset
Dataset Card
Dataset Name: Seatizen Atlas Image DatasetTask: Multi-label image classificationDomain: Marine BiodiversityLicense: cc0-1.0Size: 14,492 annotated images
Description
The Seatizen Atlas Image Dataset is a large-scale collection of annotated underwater images designed for training and evaluating artificial intelligence models in marine biodiversity research. It is specifically tailored for multi-label image… See the full description on the dataset page: https://huggingface.co/datasets/lombardata/seatizen_atlas_image_dataset.StreetView-Image-Dataset-10K
Urban Streetscape Dataset for Vision Language Models
A curated subset of 10,000 street view images with 25 essential features optimized for training vision language models on urban environment analysis tasks.
Dataset Description
This dataset contains street view imagery paired with comprehensive annotations covering infrastructure characteristics, visual perception metrics, environmental context, and semantic segmentation data.
This comprehensive dataset represents a… See the full description on the dataset page: https://huggingface.co/datasets/Sadhana-24/StreetView-Image-Dataset-10K.text-2-image-dpo-human-preferences-full
Text-2-Image DPO Human Preferences (Full)
The complete human preference dataset for text-to-image generation. 416,360 pairwise judgments from ~20,000 annotators comparing AI-generated images across two evaluation dimensions: prompt alignment and overall preference.
This is the full, unfiltered version with uniform vote weights. For quality-filtered subsets with calibrated annotator weighting, see:
datapointai/text-2-image-dpo-human-preferences (5,000 pairs, trust-weighted)… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-dpo-human-preferences-full.2026-24679-image-dataset
24-679 (Fall 2026): Sweet and Savory Food Images
kwongnon/2026-24679-image-dataset
A binary image-classification dataset of sweet and savory foods, prepared as square RGB images for
training and evaluating image-classification models. Images are organized into two classes:
0 = sweet and 1 = savory.
The intended use is a classroom machine-learning exercise focused on image preprocessing,
augmentation, transfer learning, and model evaluation rather than production food… See the full description on the dataset page: https://huggingface.co/datasets/kwongnon/2026-24679-image-dataset.2026-24679-image-dataset
24-679 (Fall 2026): Campus Recycling and Trash Bins
ccm/2026-24679-image-dataset
Classroom photographs of campus recycling and trash bins, prepared as square RGB images with four
separately generated training variants. The intended exercise is transfer learning and evaluation on
a small image collection, rather than deployment as a waste-sorting system.
Source and task
The preparation notebook uses the Google Forms export Image Data.csv and a ZIP of the form's… See the full description on the dataset page: https://huggingface.co/datasets/ccm/2026-24679-image-dataset.eurosat-dataset-with-image
EuroSAT Dataset
Overview
This dataset contains satellite images from the EuroSAT datasetThe dataset consists of RGB images with 10 different classes, each representing a distinct type of land use.
Dataset Summary
Classes: 10 (e.g., Annual Crop, Forest, Herbaceous Vegetation, Highway, Industrial, Pasture, Permanent Crop, Residential, River, Sea/Lake)
Number of Images: 27,000+ images split into training and validation sets
Image Size: 64x64 pixels, 3 channels… See the full description on the dataset page: https://huggingface.co/datasets/MuafiraThasni/eurosat-dataset-with-image.Infographic_image_dataset_samples
InfoBay.AI Infographic Dataset Catalogue
Overview
The InfoBay.AI Infographic Dataset Catalogue is a professionally curated collection of 90,000 infographic images designed for Artificial Intelligence, Computer Vision, Document Understanding, Visual Question Answering (VQA), Optical Character Recognition (OCR), and Multimodal Foundation Models.
The collection includes high-quality infographic assets covering business intelligence, education, finance, healthcare… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Infographic_image_dataset_samples.2026-24679-image-dataset
24-679 (Fall 2026): 3D-Printed vs Manufactured Household Objects
kadireks/2026-24679-image-dataset
A binary image-classification dataset of everyday objects, prepared as square RGB images for training
and evaluating image-classification models. Images are organized into two classes:
0 = not 3D printed and 1 = 3D printed.
The intended use is a classroom machine-learning exercise focused on image preprocessing,
augmentation, transfer learning, and model evaluation rather than… See the full description on the dataset page: https://huggingface.co/datasets/kadireks/2026-24679-image-dataset.2026-mug-image-dataset
Mug Classification Dataset
pcwoods/2026-mug-image-dataset
This dataset contains tabletop photographs labeled for the presence of one or more mugs. Images are prepared as 224x224 RGB squares with gray padding.
Source
All images for this dataset were hand-captured by the author and hand-labeled.
Dataset Structure
image: 224x224 RGB image.
label: 0 = No mug, 1 = Mug.
provenance: includes source_id, parent_id, and augmentation type.… See the full description on the dataset page: https://huggingface.co/datasets/pcwoods/2026-mug-image-dataset.core_sample_image_data
🖼 Soil Core Sample Image Data
This dataset contains 3604 labeled images of soil core samples for image classification.
📌 Dataset Summary
Images: Squared images (300x300 pixels), cropped from full-scale high-resolution images of soil core samples.
Labels: hb, nb (terms according to DIN 4023:2006-02).
Format: Hugging Face datasets.Dataset with Image() feature.
Split: Train / val / test split is performed with a ratio of 0.8 / 0.1 / 0.1, whereas the samples are stratified… See the full description on the dataset page: https://huggingface.co/datasets/grano1/core_sample_image_data.2025-24679-image-dataset
Dataset Card for Kaikai vs Georgie Image Dataset
Dataset Description
This dataset was created as part of a course project for 24-679.It supports binary image classification of two student-created characters, Kaikai and Georgie, to explore dataset creation, augmentation, and reproducibility workflows.
Dataset Summary
Binary classification dataset (Kaikai vs Georgie).
Student-created images with augmentations.
Educational purpose only, not intended for… See the full description on the dataset page: https://huggingface.co/datasets/cassieli226/2025-24679-image-dataset.2025-24679-image-dataset
🏞️ Nature vs. Not Nature Image Dataset
📝 Dataset Summary
This is a student-created image dataset designed for binary image classification. The dataset consists of 31 original photographs, each manually labeled as either depicting a nature scene or not_nature (e.g., man-made objects, indoor scenes).
To facilitate model training, a larger augmented split is provided, expanding the dataset to 310 images through a series of documented, label-preserving transformations. The… See the full description on the dataset page: https://huggingface.co/datasets/zacCMU/2025-24679-image-dataset.kortowo-dynamics-image-dataset
Kortowo Dynamics Image Dataset
This is the image dataset for the Kortowo Dynamics project.
It contains frames derived from the original video dataset.
Purpose
The initial purpose of this dataset has been to train a video classification model using the strategy of analysing individual frames.
Contents
The dataset contains video frames of the Boston Dynamics' Spot robot performing different actions.
Available classes:
body_swing
crawling
jumping
looking_left… See the full description on the dataset page: https://huggingface.co/datasets/jlynxdev/kortowo-dynamics-image-dataset.figma_image_dataset_samples
InfoBay.AI Design Assets Dataset Collection
Overview
The InfoBay.AI Design Assets Dataset Collection is a comprehensive multimodal dataset containing more than 2.37 million professionally curated digital design assets across multiple creative domains. The collection has been prepared for Artificial Intelligence, Machine Learning, Computer Vision, Design Automation, Multimodal Foundation Models, Enterprise Search, Retrieval Systems, UI/UX Research, and Generative… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/figma_image_dataset_samples.2025-24679-image-dataset
Dataset Card for ccm/2025-24679-image-dataset
Dataset Details
Dataset Description
This dataset consists of images labeled as recycling (0) or trash (1). It was created as part of a classroom exercise in supervised learning and data augmentation, with the goal of giving students practice in building and evaluating image classification pipelines.
Curated by: Fall 2025 24-679 course at Carnegie Mellon University
Shared by [optional]: Christopher McComb… See the full description on the dataset page: https://huggingface.co/datasets/ccm/2025-24679-image-dataset.hw1-24679-image-dataset
Asian vs Western Food Classification Dataset
Dataset Summary
Purpose: This dataset was created for binary classification of food images into Asian or Western cuisine categories, developed as part of CMU 24-679 coursework to explore computer vision techniques in food recognition.
Quick Stats:
360 total images (40 original + 320 augmented)
Binary classification task
224x224 RGB images
Balanced classes (~50% each category)
Contact: maryzhang@cmu.edu
Sample… See the full description on the dataset page: https://huggingface.co/datasets/maryzhang/hw1-24679-image-dataset.outfit-image-dataset
Dataset Card for Outfit Dataset
This dataset card documents the Clothing Outfit Dataset.It contains 30 original outfit images (shirt-pant combinations) with categorical and binary labels, with an augmented split expanding to 300 images.
Dataset Details
Dataset Description
Curated by: Bareethul Kader (Carnegie Mellon University)
Language(s): English (labels: formality, binary target)
License: CC BY 4.0
Repository: bareethul/clothing-outfits
Uses… See the full description on the dataset page: https://huggingface.co/datasets/bareethul/outfit-image-dataset.2025-24679-image-dataset-StefanovBooks-image-datasetHomework1_image_dataset
Baby Yoda LEGO Presence (224×224)
Binary image dataset indicating whether an image contains a Baby Yoda LEGO figure.
Composition & Collection
Creators: Student-captured photos by the dataset author for an academic assignment.
Subjects: Desktop/object scenes that may contain a Baby Yoda LEGO figure.
Original count: {orig_n} images (≥30 required).
Resolution / format: Center-cropped square, resized to 224×224, saved as JPEG.
Privacy: No faces or personally identifiable… See the full description on the dataset page: https://huggingface.co/datasets/kevinkyi/Homework1_image_dataset.ImageDataHW1
Dataset Card for ImageDataHW1
This dataset has image data of fencing competition, and is meant for classifying whether or not a point has been awarded
Dataset Details
Dataset Description
The original split contains thirty screenshots retrieved from the source below. 15 images of the original split show frames where a point has not yet been
awarded, while 15 show frames where a point has been awarded. The augmented split shows 300 images that have been… See the full description on the dataset page: https://huggingface.co/datasets/emkessle/ImageDataHW1.2025-24679-image-dataset
Car Classification Dataset - Original
Dataset Description
This dataset contains 31 original car images collected for binary classification tasks. Images are captured from various angles and in different lighting conditions.
Dataset Summary
Images: 31 high-quality car photographs
Resolution: 224x224 pixels
Format: RGB images (converted from HEIC/PNG)
Labels: Binary classification (sedan vs SUV)
Usage
This dataset is designed for:
Image… See the full description on the dataset page: https://huggingface.co/datasets/Anyuhhh/2025-24679-image-dataset.desdep_image_datasettext-2-image-dpo-human-preferences
Text-2-Image DPO Human Preferences
A large-scale, quality-controlled human preference dataset for text-to-image generation. 80,000 trust-weighted pairwise judgments from calibrated annotators comparing AI-generated images across two evaluation dimensions: prompt alignment and overall preference.
Built on the Datapoint annotation platform — purpose-built infrastructure for collecting high-quality human preference data at scale.
Overview
Metric
Value
Total… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-dpo-human-preferences.text-2-image-dpo-human-preferences-small
Text-2-Image DPO Human Preferences (Small)
A quality-controlled human preference dataset for text-to-image generation. 40,000 trust-weighted pairwise judgments from calibrated annotators comparing AI-generated images across two evaluation dimensions: prompt alignment and overall preference.
This is the highest-annotator-quality subset. For the full 5,000-pair dataset, see datapointai/text-2-image-dpo-human-preferences.
Built on the Datapoint annotation platform — purpose-built… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-dpo-human-preferences-small.
