datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
danbooru-multitier-captions-202606
Danbooru — multi-tier captions (202606)
Per-post Danbooru data for the 202606 crawl: native tags, the raw API metadata, model-generated
multi-tier natural-language captions (long / refined long / medium / short), and post flags.
One row per Danbooru post_id. Images are not included — each post is referenced by
post_id, danbooru_url, md5, and the Danbooru CDN URLs. (The two example previews below are
downscaled for illustration.)
Based on:… See the full description on the dataset page: https://huggingface.co/datasets/BootsofLagrangian/danbooru-multitier-captions-202606.moss-character-voices-top3-captioned
MOSS Character Voices — Top-3 per Group, Captioned (training-ready)
The top-3 takes per group (rank 0/1/2 by reward model) of all ~12,900 groups in
laion/moss-character-voices-bestof64
— ~38,700 samples — each annotated with both LAION voice-acting caption pipelines and shipped with
all scores + metadata, ready to drop into training. Sharded as data/train-*-of-00129.parquet.
Captions (two pipelines, same clip)
caption_procedural — Procedural Voice Captions:
terse… See the full description on the dataset page: https://huggingface.co/datasets/laion/moss-character-voices-top3-captioned.mscoco-1st-captionTo reproduce, run pip install -r requirements.txt and download.sh.
flickr30k_clip-SimCLRv2-caption_pairsflickr30k_clip-ViT-B-32-caption_pairslaion400m-person-captionshowto100m_captions_with_verb_nounslaion-coco-13m-molmo-d-7b
Dataset Card for laion-coco-13m-molmo-d-7b
Dataset Summary
This is 41,409,699 new synthetic captions for the 13,803,233 images found in laion/laion-coco. It includes the original captions from that repository as well as new captions. The dataset was filtered to images >= 512px on the short edge.
The long captions were produced using allenai/Molmo-7B-D-0924. Medium and short captions were produced from these captions using allenai/Llama-3.1-Tulu-3-8B-DPO. The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/laion-coco-13m-molmo-d-7b.compositionality_eccv_captionflickr-megalith-10m-internvl2-multi-caption
Dataset Card for flickr-megalith-10m-internvl2-multi-caption
Dataset Summary
This is approximately 57.3 million synthetic captions for the images found in madebyollin/megalith-10m.
It includes the following captions:
InternVL2 8B long captions (by CaptionEmporium)
InternVL2 8B short captions (by CaptionEmporium)
Florence2 long captions (by aipicasso)
Florence2 short captions (by CaptionEmporium)
ShareCaptioner long captions (by drawthingsai)
ShareCaptioner short… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/flickr-megalith-10m-internvl2-multi-caption.conceptual_captions_3m_zh_tiny_0
Dataset Card for "conceptual_captions_3m_zh_tiny_0"
More Information needed
Re-LAION-Caption19M
Re-LAION-Caption 19M
This dataset is based on the paper Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
Re-LAION-Caption 19M is a high-quality, recaptioned subset of Re-LAION-5B consisting of 19 million 1024×1024 images with structured captions. This dataset was curated to improve prompt adherence and alignment in text-to-image generative models.
Motivation
Most large-scale image-text datasets (e.g., LAION-5B) suffer from… See the full description on the dataset page: https://huggingface.co/datasets/supermodelresearch/Re-LAION-Caption19M.conceptual_captions_3m_zh_tiny_4
Dataset Card for "conceptual_captions_3m_zh_tiny_4"
More Information needed
conceptual_captions_3m_zh_tiny_2
Dataset Card for "conceptual_captions_3m_zh_tiny_2"
More Information needed
conceptual_captions_3m_zh_tiny_5
Dataset Card for "conceptual_captions_3m_zh_tiny_5"
More Information needed
conceptual_captions_3m_zh_tiny_1
Dataset Card for "conceptual_captions_3m_zh_tiny_1"
More Information needed
mscoco_train_2014_openai_clip-vit-base-patch32_image_caption_retrieval_pairs_2022-09-01conceptual_captions_3m_zh_tiny_3
Dataset Card for "conceptual_captions_3m_zh_tiny_3"
More Information needed
pick-double-caption
Dual Caption Preference Optimization for Diffusion Models
We propose DCPO, a new paradigm to improve the alignment performance of text-to-image diffusion models. For more details on the technique, please refer to our paper here.
Developed by
Amir Saeidi*
Yiran Luo*
Agneet Chatterjee
Shamanthak Hegde
Bimsara Pathiraja
Yezhou Yang
Chitta Baral
Dataset
This dataset is Pick-Double Caption, a modified version of the Pick-a-Pic V2 dataset. We… See the full description on the dataset page: https://huggingface.co/datasets/DualCPO/pick-double-caption.pexels-people-captions
Pexels people captions
37,412 photographs of people, each with a description written for it in
Chinese or English. The photographs themselves are not included: every row
carries a link to the original on Pexels instead.
What a row holds
column
meaning
image_id
stable identifier used across the corpus, pexels-<photo id>
photo_id
the Pexels photo id
page_url
the photo's page on pexels.com
image_url
the original image file as the API reports it… See the full description on the dataset page: https://huggingface.co/datasets/kaupane/pexels-people-captions.vintage-photography-450k-high-quality-captionsThis is a 450k image datastet focused on photography from the 20th century, and their analog aspect. Many of the images are in high resolution. This dataset currently has 20k images captioned with InternVL2 26B, and is a work in progress (I plan to caption the entire dataset and also have short captions for all of the images, compute is an issue for now).
ffhq_with_llava_shorter_captions_flux_latentst2i-nla-caption-50kprompt-training-data-public
T2I-NLA Caption 50k-Prompt Training Data
Public backup of final T2I-NLA reader training data.
Broader caption-data training set from the 50k-prompt caption-reader run.
This repository is intended to make final-reader retraining possible after the
original server is no longer available.
See manifest.json for the original local path, size, and associated reader models.
captioned-vqav2captioned-ChordStream
Captioned MIDI with ChordStream
1,000 captioned MIDI files, each with its harmony extracted as a ChordStream token
sequence: one 64-bit token per chord change or silence, carrying the key, the chord's
degree in that key, its pitch-class content, its bass, and its absolute bar, onset and
length in musical time.
Rows are the first 1,000 midicaps_prose rows of PsiPi/captioned-midi-moods-and-genre
(a private dataset), with the ChordStream column computed from each row's own midi… See the full description on the dataset page: https://huggingface.co/datasets/PsiPi/captioned-ChordStream.vg-captions-graphs-processed-image-graphsvintage-artworks-60k-captionedThis is a dataset consisting of 60k vintage artworks from the 20th century, consisting of vintage pulp, sci-fi and pinup artworks from that era.
The dataset has short and long captions for each image, as well as resolution information. The large captions (large_caption column) were made with florence-2-large-ft, and then shortened with llama 3 8b (see short_caption column).
PokeFA-pokemon-fanart-captioned
PokeFA — Pokémon fan-art metadata with relevance/aesthetic scores and hybrid captions
PokeFA is a large-scale Pokémon fan-art dataset released as metadata + URLs only (no image bytes).~30,000 candidate images are collected across 1,025 Pokémon using a popularity-banded budget with following curation pipeline:
NSFW filtering → OCR localization & inpainting → resizing → relevance & aesthetic scoring (GPT-5-mini vision) → near-duplicate removal → quality filtering to the top ~16… See the full description on the dataset page: https://huggingface.co/datasets/Kev0208/PokeFA-pokemon-fanart-captioned.leetcode_with_youtube_captionsdan-caption-twostage-synth-full
