datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dummy_image_text_data
Dataset Card for "dummy_image_text_data"
More Information needed
open-image-preferences-v1
Open Image Preferences
Prompt: Anime-style concept art of a Mayan Quetzalcoatl biomutant, dystopian world, vibrant colors, 4K.
Image 1
Image 2
Prompt: 8-bit pixel art of a blue knight, green car, and glacier landscape in Norway, fantasy style, colorful and detailed.
Image 1… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/open-image-preferences-v1.Defactify_Image_Dataset
Defactify_Image_Dataset
This dataset is associated with the paper A Comprehensive Dataset for Human vs. AI Generated Image Detection.
📝 Dataset Description
Dataset Summary
The Defactify_Image_Dataset (A Comprehensive Dataset for Human vs. AI Generated Image Detection) is a high-quality collection of 96,000 images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. Built using the MS… See the full description on the dataset page: https://huggingface.co/datasets/Rajarshi-Roy-research/Defactify_Image_Dataset.pagoda-text-and-image-dataset
Dataset Card for "pagoda-text-and-image-dataset"
More Information needed
yt_full_image_dataset
Dataset Card for "yt_full_image_dataset"
More Information needed
fashion-image-datasettext-2-image-human-preferences-2m
Text-to-image human preferences: 2M votes across 30 models
This dataset contains the complete voting record behind the
Datapoint Image Bench
leaderboard: 2,161,160 validated pairwise votes — exactly 10 for each of
216,116 image pairs. The votes compare 30 text-to-image models in a complete
round-robin on 500 prompts, judged by annotators from over 200 countries.
Every vote includes the annotator's trust score at the time the vote was
cast.
Built on the Datapoint annotation… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-human-preferences-2m.open-image-preferences-v1-binarized
Open Image Preferences
Prompt: Anime-style concept art of a Mayan Quetzalcoatl biomutant, dystopian world, vibrant colors, 4K.
Image 1
Image 2
Prompt: 8-bit pixel art of a blue knight, green car, and glacier landscape in Norway, fantasy style, colorful and detailed.
Image 1… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/open-image-preferences-v1-binarized.yt_main_image_dataset
Dataset Card for "yt_main_image_dataset"
More Information needed
my-image-caption-datasetimage-text-dataset-subset-300k-captions_onlygovdocs1-image
BEE-spoke-data/govdocs1-image
This contains .jpg files from govdocs1. Light deduplication was applied (i.e. jdupes on all files) which removed ~500 duplicate images.
DatasetDict({
train: Dataset({
features: ['image'],
num_rows: 108895
})
})
source
Source info/page: https://digitalcorpora.org/corpora/file-corpora/files/
@inproceedings{garfinkel2009bringing,
title={Bringing Science to Digital Forensics with Standardized Forensic Corpora}… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-image.pagoda-text-and-image-dataset-small
Dataset Card for "pagoda-text-and-image-dataset-small"
More Information needed
brain-tumor-image-dataset-semantic-segmentation
Dataset Card for "brain-tumor-image-dataset-semantic-segmentation"
Dataset Description
The Brain Tumor Image Dataset (BTID) for Semantic Segmentation contains MRI images and annotations aimed at training and evaluating segmentation models. This dataset was sourced from Kaggle and includes detailed segmentation masks indicating the presence and boundaries of brain tumors.
This dataset can be used for developing and benchmarking algorithms for medical image segmentation… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/brain-tumor-image-dataset-semantic-segmentation.shopping-queries-image-dataset
Shopping Queries Image Dataset (SQID 🦑): An Image-Enriched ESCI Dataset for Exploring Multimodal Learning in Product Search
Introduction
The Shopping Queries Image Dataset (SQID) is a dataset that includes image information for over 190,000 products. This dataset is an augmented version of the Amazon Shopping Queries Dataset, which includes a large number of product search queries from real Amazon users, along with a list of up to 40 potentially relevant results and… See the full description on the dataset page: https://huggingface.co/datasets/crossingminds/shopping-queries-image-dataset.low-alt-satellite-image-dataset-5k-sam3-segmented_jsonmy_image_captioning_datasetwebsite_screenshots_image_dataset
Website Screenshots Image Dataset
This dataset is obtainable here from roboflow..
Dataset Details
Dataset Description
Language(s) (NLP): [English]
License: [MIT]
Dataset Sources
Source: [https://universe.roboflow.com/roboflow-gw7yv/website-screenshots/dataset/1]
Uses
From the roboflow website:
Annotated screenshots are very useful in Robotic Process Automation. But they can be expensive to label. This dataset would cost over… See the full description on the dataset page: https://huggingface.co/datasets/Zexanima/website_screenshots_image_dataset.image-datasetimage_diff_data
VDiff-Bench
VDiff-Bench is a multiple-choice benchmark for fine-grained visual difference identification. Each example presents two similar images and four candidate descriptions, exactly one of which states a real difference between the images.
Dataset structure
The train split contains 1,756 questions with the following fields:
id: stable example identifier.
image_1, image_2: the paired images.
choices: an object containing answer choices A, B, C, and D.… See the full description on the dataset page: https://huggingface.co/datasets/elaine1wan/image_diff_data.low-alt-satellite-image-dataset-5k-sam3-segmentedmultimodal_qa_dataset_v2_image_focuseyes-df-image-datasetStreetView-Image-Dataset-10K-train-test-splitStreetView-Image-Dataset-10K
Urban Streetscape Dataset for Vision Language Models
A curated subset of 10,000 street view images with 25 essential features optimized for training vision language models on urban environment analysis tasks.
Dataset Description
This dataset contains street view imagery paired with comprehensive annotations covering infrastructure characteristics, visual perception metrics, environmental context, and semantic segmentation data.
This comprehensive dataset represents a… See the full description on the dataset page: https://huggingface.co/datasets/Sadhana-24/StreetView-Image-Dataset-10K.text-2-image-dpo-human-preferences-full
Text-2-Image DPO Human Preferences (Full)
The complete human preference dataset for text-to-image generation. 416,360 pairwise judgments from ~20,000 annotators comparing AI-generated images across two evaluation dimensions: prompt alignment and overall preference.
This is the full, unfiltered version with uniform vote weights. For quality-filtered subsets with calibrated annotator weighting, see:
datapointai/text-2-image-dpo-human-preferences (5,000 pairs, trust-weighted)… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-dpo-human-preferences-full.nepali-image-caption-datasetnew-masked-jaw-df-image-datasetmouth-df-image-datasetnew-masked-forehead-df-image-dataset
