datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text-to-image-promptsIf you have questions about this dataset , feel free to ask them on the fusion-discord : https://discord.gg/8TVHPf6Edn
This collection contains sets from the fusion-t2i-ai-generator on perchance.
This datset is used in this notebook: https://huggingface.co/datasets/codeShare/text-to-image-prompts/tree/main/Google%20Colab%20Notebooks
To see the full sets, please use the url "https://perchance.org/" + url
, where the urls are listed below:
_generator
gen_e621
fusion-t2i-e621-tags-1… See the full description on the dataset page: https://huggingface.co/datasets/codeShare/text-to-image-prompts.Describable-Textures-Dataset
Dataset Card for Describable Textures Dataset
This is a FiftyOne dataset with 5640 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
import fiftyone.utils.huggingface as fouh
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = fouh.load_from_hub("Voxel51/Describable-Textures-Dataset")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Describable-Textures-Dataset.TextRich
Dataset Composition
This dataset is a multi-domain bennchmark for detecting AI-generated text-rich images from GPT-Image-2.
The dataset consists of two complementary subsets:
Fake subset:
Images generated by GPT-Image-2 using carefully designed prompts. The prompts are constructed to cover diverse domains and layouts while avoiding reproduction of specific real-world images.
Real subset:
Images sampled from six publicly available datasets. These images are selected as… See the full description on the dataset page: https://huggingface.co/datasets/Shuyiww/TextRich.text-2-image-Rich-Human-Feedback
Building upon Google's research Rich Human Feedback for Text-to-Image Generation we have collected over 1.5 million responses from 152'684 individual humans using Rapidata via the Python API. Collection took roughly 5 days.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
We asked humans to evaluate AI-generated images in style, coherence and prompt alignment. For images that contained flaws, participants were… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-image-Rich-Human-Feedback.text-to-image-2M
text-to-image-2M: A High-Quality, Diverse Text-to-Image Training Dataset
Citation
@article{zou2026advancing,
title = {Advancing Aesthetic Image Generation via Composition Transfer},
author = {Zou, Kai and Zhao, Zhiwei and Liu, Bin and Yu, Nenghai},
journal = {International Journal of Computer Vision},
volume = {134},
pages = {252},
year = {2026},
doi = {10.1007/s11263-026-02862-8},
url = {https://doi.org/10.1007/s11263-026-02862-8}… See the full description on the dataset page: https://huggingface.co/datasets/jackyhate/text-to-image-2M.text-2-image-human-preferences-2m
Text-to-image human preferences: 2M votes across 30 models
This dataset contains the complete voting record behind the
Datapoint Image Bench
leaderboard: 2,161,160 validated pairwise votes — exactly 10 for each of
216,116 image pairs. The votes compare 30 text-to-image models in a complete
round-robin on 500 prompts, judged by annotators from over 200 countries.
Every vote includes the annotator's trust score at the time the vote was
cast.
Built on the Datapoint annotation… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-human-preferences-2m.Textual-Image-Caption-Dataset
Update: OCT-2023
Add v2 with recent SoTA model swinV2 classifier for both soft/hard-label visual_caption_cosine_score_v2 with person label (0.2, 0.3 and 0.4)
Introduction
Modern image captaining relies heavily on extracting knowledge, from images such as objects,
to capture the concept of static story in the image. In this paper, we propose a textual visual context dataset
for captioning, where the publicly available dataset COCO caption (Lin et al., 2014) has been… See the full description on the dataset page: https://huggingface.co/datasets/AhmedSSabir/Textual-Image-Caption-Dataset.cc0-textures
Dataset Card for CC0 Textures
Dataset Summary
This dataset contains 18,785 texture images from cc0-textures.com. It includes textures of wood, metal, concrete, fabric, stone, ceramic, and other materials. The original archives were downloaded, unpacked, and images were compressed using PNG optimization and JPEG quality compression (90%) to reduce file size while keeping good quality.
Languages
The dataset is monolingual:
English (en): Texture titles and tags… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/cc0-textures.NFT-70M_text
Dataset Card for "NFT-70M_text"
Dataset summary
The NFT-70M_text dataset is a companion for our released NFT-70M_transactions dataset,
which is the largest and most up-to-date collection of Non-Fungible Tokens (NFT) transactions between 2021 and 2023 sourced from OpenSea.
As we also reported in the "Data anonymization" section of the dataset card of NFT-70M_transactions,
the textual contents associated with the NFT data were replaced by identifiers to numerical… See the full description on the dataset page: https://huggingface.co/datasets/MLNTeam-Unical/NFT-70M_text.Kather-texture-2016
Collection of textures in colorectal cancer histology
Description
This data set represents a collection of textures in histological images of human colorectal cancer.
It contains 5000 histological images of 150 * 150 px each (74 * 74 µm). Each image belongs to exactly one of eight tissue categories.
Image format
All images are RGB, 0.495 µm per pixel, digitized with an Aperio ScanScope (Aperio/Leica biosystems), magnification 20x.
Histological samples are… See the full description on the dataset page: https://huggingface.co/datasets/1aurent/Kather-texture-2016.text-2-image-dpo-human-preferences-full
Text-2-Image DPO Human Preferences (Full)
The complete human preference dataset for text-to-image generation. 416,360 pairwise judgments from ~20,000 annotators comparing AI-generated images across two evaluation dimensions: prompt alignment and overall preference.
This is the full, unfiltered version with uniform vote weights. For quality-filtered subsets with calibrated annotator weighting, see:
datapointai/text-2-image-dpo-human-preferences (5,000 pairs, trust-weighted)… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-dpo-human-preferences-full.TextRich
Dataset Composition
This dataset is a multi-domain bennchmark for detecting AI-generated text-rich images from GPT-Image-2.
The dataset consists of two complementary subsets:
Fake subset:
Images generated by GPT-Image-2 using carefully designed prompts. The prompts are constructed to cover diverse domains and layouts while avoiding reproduction of specific real-world images.
Real subset:
Images sampled from six publicly available datasets. These images are selected as… See the full description on the dataset page: https://huggingface.co/datasets/Luc7d/TextRich.textures3
Description
This is the third iteration and official release of a dataset curated to power the Materializer model for Blender. The dataset contains a range of labeled texture images that were sourced from ambientCG under their Creative Commons CC0 1.0 Universal License. These textures are designed to help in the classification of various material maps, which are essential for creating realistic 3D materials in Blender.
Future Plans
The dataset is still evolving, and I… See the full description on the dataset page: https://huggingface.co/datasets/DeathDaDev/textures3.texture-shape-cue-conflictThis dataset contains the stimuli for ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness by Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel.
The stimuli allow testing of the texture/shape bias in an observer model (human or artificial) by containing two conflicting cues per image (shape and texture). The images generated using iterative style transfer (Gatys et al., 2016) between… See the full description on the dataset page: https://huggingface.co/datasets/rgeirhos/texture-shape-cue-conflict.text-to-image-2M
text-to-image-2M: A High-Quality, Diverse Text-to-Image Training Dataset
Overview
text-to-image-2M is a curated text-image pair dataset designed for fine-tuning text-to-image models. The dataset consists of approximately 2 million samples, carefully selected and enhanced to meet the high demands of text-to-image model training. The motivation behind creating this dataset stems from the observation that datasets with over 1 million samples tend to produce better… See the full description on the dataset page: https://huggingface.co/datasets/NovaBrown/text-to-image-2M.textureninja
Dataset Card for Texture Ninja
Dataset Summary
This dataset contains 4,540 texture images from texture.ninja. It includes high-resolution textures of brick, concrete, rock, wood, metal, paint, plaster, ground materials, and other surfaces. The original images were downloaded, processed, and compressed using PNG optimization and JPEG quality compression (90%) to reduce file size while maintaining good quality.
Languages
The dataset is monolingual:
English (en):… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/textureninja.text-to-image-2M
text-to-image-2M: A High-Quality, Diverse Text-to-Image Training Dataset
Overview
text-to-image-2M is a curated text-image pair dataset designed for fine-tuning text-to-image models. The dataset consists of approximately 2 million samples, carefully selected and enhanced to meet the high demands of text-to-image model training. The motivation behind creating this dataset stems from the observation that datasets with over 1 million samples tend to produce better… See the full description on the dataset page: https://huggingface.co/datasets/rivisia/text-to-image-2M.text-2-image-dpo-human-preferences
Text-2-Image DPO Human Preferences
A large-scale, quality-controlled human preference dataset for text-to-image generation. 80,000 trust-weighted pairwise judgments from calibrated annotators comparing AI-generated images across two evaluation dimensions: prompt alignment and overall preference.
Built on the Datapoint annotation platform — purpose-built infrastructure for collecting high-quality human preference data at scale.
Overview
Metric
Value
Total… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-dpo-human-preferences.Kather-texture-2016
Collection of textures in colorectal cancer histology
Description
This data set represents a collection of textures in histological images of human colorectal cancer.
It contains 5000 histological images of 150 * 150 px each (74 * 74 µm). Each image belongs to exactly one of eight tissue categories.
Image format
All images are RGB, 0.495 µm per pixel, digitized with an Aperio ScanScope (Aperio/Leica biosystems), magnification 20x.
Histological samples are… See the full description on the dataset page: https://huggingface.co/datasets/YueFanXia/Kather-texture-2016.texturecan
Dataset Card for TextureCan Textures
Dataset Summary
This dataset contains 4,037 texture images from texturecan.com. It includes textures of various materials such as brick, paper, fabric, metal, wood, stone, and other surfaces. The original archives were downloaded, unpacked, and images were compressed using PNG optimization and JPEG quality compression (90%) to reduce file size while maintaining good quality.
Languages
The dataset is monolingual:
English… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/texturecan.socratis_image_text_emotion
SOCRATIS: A benchmark of diverse open-ended emotional reactions to image-caption pairs.
ICCV WECIA Workshop 2023 (oral)
Project Page, Paper
We release a benchmark which contains 18K diverse emotions and reasons for feeling them on 2K image-caption pairs.
Our current preliminary findings have shown that Humans prefer human-written emotional reactions over machine-generated by more than two times.
We also find that current metrics fail to correlate with human preference… See the full description on the dataset page: https://huggingface.co/datasets/array/socratis_image_text_emotion.Technical-Textile-Synthesis-Luxury-Lingerie-Embroidery
🩰 BWS La Perla Premium: Compliance-Native Multimodal Tokens (POC)
🛡️ Engineering Evaluation Sandbox (Active 7-Day Access)
Technical Ingestion Portal: s3://createphotos (Whitelisted buckets only)
Secure Evaluation Link: Download 03_La_Perla__lingerie_30_Enterprise_POC.zip
Direct Manifest Auditor: BWS Forensic Manifest Repository
Procurement: All assets are 2026 US CLEAR Act compliant. Access is granted to whitelisted engineering nodes only. Forward your AWS Account ID to… See the full description on the dataset page: https://huggingface.co/datasets/BWS-Data-Solutions/Technical-Textile-Synthesis-Luxury-Lingerie-Embroidery.text-to-image-promptsIf you have questions about this dataset , feel free to ask them on the fusion-discord : https://discord.gg/8TVHPf6Edn
This collection contains sets from the fusion-t2i-ai-generator on perchance.
This datset is used in this notebook: https://huggingface.co/datasets/codeShare/text-to-image-prompts/tree/main/Google%20Colab%20Notebooks
To see the full sets, please use the url "https://perchance.org/" + url
, where the urls are listed below:
_generator
gen_e621
fusion-t2i-e621-tags-1… See the full description on the dataset page: https://huggingface.co/datasets/bofbofai/text-to-image-prompts.text2glioma-synthetic-10k
Text2Glioma synthetic MRI dataset (10k)
10,000 synthetic 4-sequence (T1, T1CE, T2, FLAIR) 3D brain MRIs at
160 × 224 × 160, 1 mm isotropic (NIfTI, .nii.gz), generated by the
Text2Glioma latent diffusion model conditioned on VASARI radiology
prompts and expert-derived tumor segmentation masks.
Clinical and commercial usage
This dataset and the Text2Glioma model are for research purposes only.
The authors expressly disavow any work that uses this dataset for… See the full description on the dataset page: https://huggingface.co/datasets/vasileionromaioi/text2glioma-synthetic-10k.text-to-image-2M
text-to-image-2M: A High-Quality, Diverse Text-to-Image Training Dataset
Overview
text-to-image-2M is a curated text-image pair dataset designed for fine-tuning text-to-image models. The dataset consists of approximately 2 million samples, carefully selected and enhanced to meet the high demands of text-to-image model training. The motivation behind creating this dataset stems from the observation that datasets with over 1 million samples tend to produce better… See the full description on the dataset page: https://huggingface.co/datasets/Junaid7188/text-to-image-2M.text-2-image-dpo-human-preferences-small
Text-2-Image DPO Human Preferences (Small)
A quality-controlled human preference dataset for text-to-image generation. 40,000 trust-weighted pairwise judgments from calibrated annotators comparing AI-generated images across two evaluation dimensions: prompt alignment and overall preference.
This is the highest-annotator-quality subset. For the full 5,000-pair dataset, see datapointai/text-2-image-dpo-human-preferences.
Built on the Datapoint annotation platform — purpose-built… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-dpo-human-preferences-small.
