datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text-to-image-promptsIf you have questions about this dataset , feel free to ask them on the fusion-discord : https://discord.gg/8TVHPf6Edn
This collection contains sets from the fusion-t2i-ai-generator on perchance.
This datset is used in this notebook: https://huggingface.co/datasets/codeShare/text-to-image-prompts/tree/main/Google%20Colab%20Notebooks
To see the full sets, please use the url "https://perchance.org/" + url
, where the urls are listed below:
_generator
gen_e621
fusion-t2i-e621-tags-1… See the full description on the dataset page: https://huggingface.co/datasets/codeShare/text-to-image-prompts.dummy_image_text_data
Dataset Card for "dummy_image_text_data"
More Information needed
Multimodal-Dataset-Image_Text_Table_TimeSeries-for-Financial-Time-Series-ForecastingThe sp500stock_data_description.csv file provides detailed information on the existence of four modalities (text, image, time series, and table) for 4,213 S&P 500 stocks.
The hs300stock_data_description.csv file provides detailed information on the existence of four modalities (text, image, time series, and table) for 858 HS 300 stocks.
If you find our research helpful, please cite our paper:
@article{xu2025finmultitime,
title={FinMultiTime: A Four-Modal Bilingual Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/Wenyan0110/Multimodal-Dataset-Image_Text_Table_TimeSeries-for-Financial-Time-Series-Forecasting.mscoco_2014_5k_test_image_text_retrieval
MSCOCO (5K test set)
Original paper: Microsoft COCO: Common Objects in Context
Homepage: https://cocodataset.org/#home
5K test set split from: http://cs.stanford.edu/people/karpathy/deepimagesent/caption_datasets.zip
Bibtex:
@inproceedings{lin2014microsoft,
title={Microsoft coco: Common objects in context},
author={Lin, Tsung-Yi and Maire, Michael and Belongie, Serge and Hays, James and Perona, Pietro and Ramanan, Deva and Doll{\'a}r, Piotr and Zitnick, C Lawrence}… See the full description on the dataset page: https://huggingface.co/datasets/nlphuji/mscoco_2014_5k_test_image_text_retrieval.flickr_1k_test_image_text_retrieval
Flickr30k (1K test set)
Original paper: From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Homepage: https://shannon.cs.illinois.edu/DenotationGraph/
1K test set split from: http://cs.stanford.edu/people/karpathy/deepimagesent/caption_datasets.zip
Bibtex:
@article{young2014image,
title={From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions}… See the full description on the dataset page: https://huggingface.co/datasets/nlphuji/flickr_1k_test_image_text_retrieval.text-2-image-Rich-Human-Feedback
Building upon Google's research Rich Human Feedback for Text-to-Image Generation we have collected over 1.5 million responses from 152'684 individual humans using Rapidata via the Python API. Collection took roughly 5 days.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
We asked humans to evaluate AI-generated images in style, coherence and prompt alignment. For images that contained flaws, participants were… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-image-Rich-Human-Feedback.image-text_medieval-scripts_xiv-xv-xvi
Dataset Card for image-text_medieval-scripts_xiv-xv-xvi
This dataset was created using pagexml-hf converter from Transkribus PageXML data.
Dataset Summary
This dataset contains 548322 samples across 1 split(s).
Geographical scope: BelgiumPeriod: 1350-1550Languages: FlemishType of document: ProtocolProvenance: State Archives in Leuven
Projects Included
Itinera Nova
Parts of Charters from Königsfelden
SAL7304_full
SAL7305_full
SAL7306_full
SAL7307
SAL7307_full… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_medieval-scripts_xiv-xv-xvi.dhivehi-image-text
Dhivehi Image-Text Dataset
A dataset of Dhivehi (Maldivian) image-text pairs for machine learning and computer vision tasks.
Dataset Statistics
Total number of batches: 10
Total images across all batches: 394,212
Average images per batch: ~39,421
Split ratios:
Training: 80%
Validation: 10%
Test: 10%
Batch Details
Batch
Total Images
Train
Validation
Test
dv01-01
39659
31727
3966
3966
dv01-02
38989
31191
3899
3899
dv01-03
39360
31488
3936
3936… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-image-text.text-to-image-2M
text-to-image-2M: A High-Quality, Diverse Text-to-Image Training Dataset
Citation
@article{zou2026advancing,
title = {Advancing Aesthetic Image Generation via Composition Transfer},
author = {Zou, Kai and Zhao, Zhiwei and Liu, Bin and Yu, Nenghai},
journal = {International Journal of Computer Vision},
volume = {134},
pages = {252},
year = {2026},
doi = {10.1007/s11263-026-02862-8},
url = {https://doi.org/10.1007/s11263-026-02862-8}… See the full description on the dataset page: https://huggingface.co/datasets/jackyhate/text-to-image-2M.surya-ocr-500-image-to-textjapanese-text-image-retrieval-trainshunk031/JDocQAのtrain splitに含まれるPDFデータを画像化し、NDLOCRでOCRしたテキストとペアにしたデータセットです。OCRは長い辺を1200pxにリサイズした画像に対して実施しました。OCR結果には、読み取りに失敗した際の文字列「〓」が含まれます。本データセットに含めている画像は、長い辺を896px、700px、588pxのいずれかにリサイズしています。どのサイズとするかは主にページに含まれる文字数で決めました。
query列は、OCR結果の文字列に対しQwen/Qwen2.5-14B-Instructで生成したものです。3つの質問を生成させ、ランダムに1つを選んだものをデータセットに含めました。質問を生成する際は以下のプロンプトを使用しました。
あなたは、質問から画像をretrieveするためのモデルをトレーニングするための(質問, 画像)ペアのデータセットを作成するプロジェクトのメンバーである。
プロジェクトは以下のように進める。
step1. ドキュメントPDFを1ページ1枚の画像ファイルに変換する
step2.… See the full description on the dataset page: https://huggingface.co/datasets/oshizo/japanese-text-image-retrieval-train.image-text_historisches-grundbuch-basel_xix-xx
Dataset Card for image-text_historisches-grundbuch-basel_xix-xx
This dataset was created using pagexml-hf converter from Transkribus PageXML data.
Dataset Summary
This dataset contains 193.409 samples across 1 split(s). The data are transcriptions (automatically generated) from volume 1 of the Historisches Grundbuch of the city of Basel. The entire collection consists of 193.409 pages, of which 135.763 pages are transcribed. The collection with Ground Truth transcriptions… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_historisches-grundbuch-basel_xix-xx.Argimi-Ardian-Finance-10k-text-image
The ArGiMI Ardian datasets : text and images
The ArGiMi project is committed to open-source principles and data sharing.
Thanks to our generous partners, we are releasing several valuable datasets to the public.
Dataset description
This dataset comprises 34,000 financial annual reports, written in English, meticulously
extracted from their original PDF format to provide a valuable resource for researchers and developers in financial
analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.unsplash-image-textThis is a dataset that streams photos data from the Unsplash 25K servers.image-text_koenigsfelden-charters-post-1500
Dataset Card for image-text_koenigsfelden-charters-post-1500
This dataset was created using pagexml-hf converter from Transkribus PageXML data.
Dataset Summary
This dataset contains 3222 samples across 1 split(s).
Geographical scope: SwitzerlandPeriod: 1291-1550Languages: Middle High German, LatinType of document: DocumentsProvenance: State Archives Aargau
Projects Included
FRAD068_03G_SAINT_PIERRE_SAINT_GILLES_032_01… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_koenigsfelden-charters-post-1500.pagoda-text-and-image-dataset
Dataset Card for "pagoda-text-and-image-dataset"
More Information needed
image-text_kurrent-xix
Dataset Card for image-text_kurrent-xix
This dataset was created using pagexml-hf converter from Transkribus PageXML data.
Dataset Summary
This dataset contains 158525 samples across 1 split(s).
Projects Included
MM_1_001
MM_1_002
MM_1_003
MM_1_004
MM_1_005
MM_1_006
MM_1_007
MM_1_008
MM_1_009
MM_1_010
MM_1_011
MM_1_012
TEST_CITlab_Bassermann_0_4
TEST_CITlab_Bassermann_Manuscripts
TEST_CITlab_Bassermann_Manuscripts_0_2
TEST_CITlab_Binder_Kochbuch_2… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_kurrent-xix.text-2-image-human-preferences-2m
Text-to-image human preferences: 2M votes across 30 models
This dataset contains the complete voting record behind the
Datapoint Image Bench
leaderboard: 2,161,160 validated pairwise votes — exactly 10 for each of
216,116 image pairs. The votes compare 30 text-to-image models in a complete
round-robin on 500 prompts, judged by annotators from over 200 countries.
Every vote includes the annotator's trust score at the time the vote was
cast.
Built on the Datapoint annotation… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-human-preferences-2m.surya-ocr-1K-image-to-texttext-2-image-Rich-Human-Feedback-32k
Building upon Google's research Rich Human Feedback for Text-to-Image Generation, and the
smaller, previous version of this dataset, we have collected over 3.7 million responses from 307'415 individual humans for the open-image-preference-v1 dataset using Rapidata via the Python API. Collection took less than 2 weeks.
If you get value from this dataset and would like to see more in the future, please consider liking it ♥️
Overview
We asked humans to evaluate AI-generated images… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-image-Rich-Human-Feedback-32k.Text_to_Image
Dataset Card
Dataset in ImagenHub.
Citation
Please kindly cite our paper if you use our code, data, models or results:
@article{ku2023imagenhub,
title={ImagenHub: Standardizing the evaluation of conditional image generation models},
author={Max Ku and Tianle Li and Kai Zhang and Yujie Lu and Xingyu Fu and Wenwen Zhuang and Wenhu Chen},
journal={arXiv preprint arXiv:2310.01596},
year={2023}
}
Text_Guided_Image_Editing
Dataset Card
Dataset in ImagenHub.
Citation
Please kindly cite our paper if you use our code, data, models or results:
@article{ku2023imagenhub,
title={ImagenHub: Standardizing the evaluation of conditional image generation models},
author={Max Ku and Tianle Li and Kai Zhang and Yujie Lu and Xingyu Fu and Wenwen Zhuang and Wenhu Chen},
journal={arXiv preprint arXiv:2310.01596},
year={2023}
}
image-text_zh-regierungsratsprotokolle
Dataset Card for image-text_zh-regierungsratsprotokolle
This dataset was created using pagexml-hf converter from Transkribus PageXML data.
Dataset Summary
This dataset contains 152786 images.
The document includes images and transcriptions of Regierungsratsprotokolle, transcribed within the frame of a project by the State Archives of Zurich.
For more information see: (https://www.zentraleserien.zh.ch/home)
Geographical scope: SwitzerlandPeriod: 1803-1883Languages:… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_zh-regierungsratsprotokolle.text-to-image-prompts
The dataset of the most popular text-to-image prompts.
Dataset Details
Dataset Description
Curated by: kazimir.ai
Funded by [optional]: [More Information Needed]
Shared by [optional]: https://kazimir.ai
License: apache-2.0
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Free to use.
Dataset Structure
CSV file… See the full description on the dataset page: https://huggingface.co/datasets/Kazimir-ai/text-to-image-prompts.Telugu-text-imageimage-text_handwritten-bundesratsprotokolle_xix-xx!!!This data set does not contain Ground Truth!!!
--- Data has been automatically created, using ATR models ---
Dataset Card for transkribus-exports-74823-raw-xml
This dataset was created using pagexml-hf converter from Transkribus PageXML data.
Dataset Summary
This dataset contains 148494 samples across 1 split(s).
These are ''automatically'' transcribed pages.
Images provided by the Federal Archives (Schweizerisches Bundesarchiv, BAR)
For a description of the project… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_handwritten-bundesratsprotokolle_xix-xx.Multimodal-Dataset-Image_Text_Table_TimeSeries-for-Financial-Time-Series-ForecastingThe sp500stock_data_description.csv file provides detailed information on the existence of four modalities (text, image, time series, and table) for 4,213 S&P 500 stocks.
The hs300stock_data_description.csv file provides detailed information on the existence of four modalities (text, image, time series, and table) for 858 HS 300 stocks.
If you find our research helpful, please cite our paper:
@article{xu2025finmultitime,
title={FinMultiTime: A Four-Modal Bilingual Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/Y123-wed/Multimodal-Dataset-Image_Text_Table_TimeSeries-for-Financial-Time-Series-Forecasting.pagoda-text-and-image-dataset-small
Dataset Card for "pagoda-text-and-image-dataset-small"
More Information needed
text-to-image-diffusiondb-2M
DiffusionDB text-to-image subset
A cleaned, safety-filtered image-prompt dataset for training a text-to-image
model, built from DiffusionDB.
Built on Hugging Face Jobs directly from poloclub/diffusiondb. It covers
part_id 1-20 (20,000 source images) before filtering. The same content is
also kept on the 20k-subset branch.
Load it with:
load_dataset("whosouravsharma/text-to-image-diffusiondb-2M")
Note on the repo name: despite "2M" in the name, this is a small slice of… See the full description on the dataset page: https://huggingface.co/datasets/whosouravsharma/text-to-image-diffusiondb-2M.image-text-dataset-subset-300k-captions_only
