datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
T2I-CoReBench-Images
T2I-CoReBench-Images
📖 Overview
T2I-CoReBench-Images is the companion image dataset of T2I-CoReBench. It contains images generated using 1,080 challenging prompts, covering both composition and reasoning scenarios undere real-world complexities.
This dataset is designed to evaluate how well current Text-to-Image (T2I) models can not only paint (produce visually consistent outputs) but also think (perform reasoning over causal chains, object relations, and logical… See the full description on the dataset page: https://huggingface.co/datasets/lioooox/T2I-CoReBench-Images.gs-images-v319c_newspapers_images_altogs-images-v2british-library-book-images
British Library Book Images
1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published
between c. 1510 and c. 1900, digitised by the British Library in partnership
with Microsoft and released by British Library Labs
on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography,
philosophy, history, poetry and literature, in several languages.
The four image types
British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/biglam/british-library-book-images.HSTLI_A-Dataset-of-Human-Semen-Time-Lapse-Images
HSTLI: A Dataset of Human Semen Time Lapse Images
Dataset Details
HSTLI contains 3,266 time-lapse microscopy videos of human sperm.Clips were recorded from two imaging modalities:
CASA system (Sperm Class Analyzer)
Optical microscope (Swift M10DB-MP + Fujifilm X-T30)
A subset of videos was manually annotated with bounding boxes around each visible sperm head.
The dataset supports detection, tracking and motility computation.
Total contents:
34… See the full description on the dataset page: https://huggingface.co/datasets/DFL-KamLab/HSTLI_A-Dataset-of-Human-Semen-Time-Lapse-Images.british-library-book-images
British Library Book Images
1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published
between c. 1510 and c. 1900, digitised by the British Library in partnership
with Microsoft and released by British Library Labs
on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography,
philosophy, history, poetry and literature, in several languages.
The four image types
British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/Faizaniqbal/british-library-book-images.dummy-base64-imageskamuicode-i2i-images
AI画像編集モデル 総合ベンチマーク v5.4
概要
本ディレクトリは、AI画像編集モデルの性能を総合的に評価するためのベンチマークスイートです。
47種類のテストを4つの難度レベルに分類し、65種類のモデル(旧45モデル + 新20モデル)を比較評価しています。
評価カバレッジ
項目
旧モデル群 (35)
新モデル群 (20)
合計
モデル数
45 (I2I/R2I 22 + T2I 13 + その他 10)
20
65
テスト数
34
47
47
生成画像
3,446
1,301 / 2,403
4,747+
VLM評価済み
完了
1,301 (54%)
進行中
評価方式
VLM-as-a-Judge: Gemini 2.5 Flash による5軸評価(Structure / Identity / Reasoning / Instruction / Quality)
スタイル多様性: photo, anime_flat, anime_cg… See the full description on the dataset page: https://huggingface.co/datasets/yumenojmd/kamuicode-i2i-images.open-imagesgenerated-imagesstanford-products-imageschartqa_without_images
Dataset Card for "chartqa_without_images"
If you wanna load the dataset, you can run the following code:
from datasets import load_dataset
data = load_dataset('ahmed-masry/chartqa_without_images')
The dataset has the following structure:
DatasetDict({
train: Dataset({
features: ['imgname', 'query', 'label', 'type'],
num_rows: 28299
})
val: Dataset({
features: ['imgname', 'query', 'label', 'type'],
num_rows: 1920
})
test:… See the full description on the dataset page: https://huggingface.co/datasets/ahmed-masry/chartqa_without_images.bbc_images_alltime
RealTimeData Monthly Collection - BBC News Images
This datasets contains all news articles head images from BBC News that were created every months from 2017 to current.
To access articles in a specific month, simple run the following:
ds = datasets.load_dataset('RealTimeData/bbc_images_alltime', '2020-02')
This will give you all BBC news head images that were created in 2020-02.
Want to crawl the data by your own?
Please head to LatestEval for the crawler… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/bbc_images_alltime.fashion-product-images-small
Dataset Card for "fashion-product-images-small"
More Information needed
Data was obtained from here
AToMiC-Images-v0.2
Dataset Card for "AToMiC-All-Images_wi-pixels"
Languages
The dataset contains 108 languages in Wikipedia.
Data Instances
Each instance is an image, its representation in bytes, and its associated captions.
Intended Usage
Image collection for Text-to-Image retrieval
Image--Caption Retrieval/Generation/Translation
Licensing Information
CC BY-SA 4.0 international license
Citation Information
TBA
Acknowledgement
Thanks… See the full description on the dataset page: https://huggingface.co/datasets/TREC-AToMiC/AToMiC-Images-v0.2.PANEL_IMAGEStelugu-synthetic-line-imagespixmo_images
PixMo Images
The raw images backing the PixMo
datasets used to train Molmo, packaged
as Parquet shards with embedded image bytes so they can be browsed in the dataset viewer
and loaded directly with datasets.
The PixMo annotation datasets (allenai/pixmo-*) ship image_urls rather than image
bytes. This repository is a content cache of those images, keyed by the SHA-256 of the
source URL.
Contents
1,073,189 images across 525 Parquet shards (data/train-*.parquet)… See the full description on the dataset page: https://huggingface.co/datasets/UWGZQ/pixmo_images.popsign-images
PopSign Images Dataset
This dataset contains frame sequences extracted from PopSign ASL (American Sign Language) video clips, organized for sign language recognition tasks.
Dataset Description
The PopSign dataset consists of short video clips of isolated ASL signs. This version provides pre-extracted image frames from each video clip, suitable for training image-based or video-based models for sign language recognition.
Subsets
The dataset contains two subsets:… See the full description on the dataset page: https://huggingface.co/datasets/sign/popsign-images.Low-Poly-Game-Asset-Images
Low-Poly Game Asset Image Dataset
A synthetic image dataset of low-poly 3D game assets with captions, built for
fine-tuning text-to-image models (LoRA / full fine-tune) on the low-poly asset domain.
Structure
dataset/
images/
p00001.png image
p00001.txt full prompt (caption)
p00001.tag.txt short object-name tag (e.g. "pistol", "tree")
...
examples/
samples_100.png preview sheet (100 samples) used in this card… See the full description on the dataset page: https://huggingface.co/datasets/laym0nd/Low-Poly-Game-Asset-Images.human-coherence-preferences-images
Rapidata Image Generation Coherence Dataset
This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
One of the largest human annotated coherence datasets for text-to-image models, this release contains over 1,200,000 human… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-coherence-preferences-images.dhivehi-vrd-images
Dhivehi Single-Line Text-Image Dataset
A collection of synthetic Dhivehi text images for training and evaluating text-image / vision models etc. Each image contains a single line of Dhivehi text with various visual styles and augmentations (Check the config field for more info on the row).
Batch Statistics
Batch
Total Images
Train
Validation
Test
vrd-batch-1
474169
379335
47417
47417
vrd-batch-2
474493
379594
47449
47450
vrd-batch-3
475564
380451
47556… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-vrd-images.pixmo-cap-images
PixMo-Cap
Big thanks to Ai2 for releasing the original PixMo-Cap dataset. To preserve the images and simplify usage of the dataset, we are releasing this version, which includes downloaded images.
PixMo-Cap is a dataset of very long (roughly 200 words on average), detailed captions.
It can be used to pre-train and fine-tune vision-language models.
PixMo-Cap was created by recording annotators speaking about an image for 60-90 seconds and then using the Claude large language model… See the full description on the dataset page: https://huggingface.co/datasets/anthracite-org/pixmo-cap-images.sdxl_images_easy_prompts-artists-seed1laion_improved_aesthetics_6.5plus_with_imagesmint-1t-html-images-gte6-sample
Size: 6769158 images sampled from Mint-1t-html
Criteria: Data entries with greater than or equal to 6 images (gte6)
touhou-imageshuman-alignment-preferences-images
Rapidata Image Generation Alignment Dataset
This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
One of the largest human annotated alignment datasets for text-to-image models, this release contains over 1,200,000 human… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-alignment-preferences-images.pixmo-cap-images
