datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
danbooru_2025_recaption
内部暂存数据集 (Internal Temporary Dataset)
English
This is a temporary dataset for internal use.
It might contain:
Items being re-processed or corrected (e.g., some images requiring re-tagging using a distributed cluster, needing a convenient data source for it).
Data to supplement our internal systems (e.g., if a machine accidentally lost some images and we don't want to re-download everything).
Recent updates or experimental data not yet finalized (e.g., the image source… See the full description on the dataset page: https://huggingface.co/datasets/NebulaeWis/danbooru_2025_recaption.JourneyDB-recaption
JourneyDB Recaption
Recaptioned version of the JourneyDB dataset using Qwen vision-language models.
Dataset Description
JourneyDB is a large-scale dataset of AI-generated images from Midjourney. This recaptioned version provides detailed visual descriptions generated by a vision-language model, which are more accurate than the original generation prompts for describing actual image content.
Statistics
Metric
Count
Total rows
3,389,605
File size… See the full description on the dataset page: https://huggingface.co/datasets/undefined443/JourneyDB-recaption.JourneyDB-recaption
JourneyDB Recaption
Recaptioned version of the JourneyDB dataset using Qwen vision-language models.
Dataset Description
JourneyDB is a large-scale dataset of AI-generated images from Midjourney. This recaptioned version provides detailed visual descriptions generated by a vision-language model, which are more accurate than the original generation prompts for describing actual image content.
Statistics
Metric
Count
Total rows
3,389,605… See the full description on the dataset page: https://huggingface.co/datasets/toilaluan/JourneyDB-recaption.human-recaption
Dataset Card for Human Recaption
This dataset contains 240,146 recaptioned images focusing on human subjects, derived from the HumanCaption-HQ-311K dataset. It provides high-quality bilingual (English and Chinese) captions, aesthetic scores, and other metadata generated using the Qwen2-VL model.
This dataset is a recaptioned version of OpenFace-CQUPT/HumanCaption-HQ-311K. The original dataset contained 313,482 samples. This version contains 240,146 samples; the reduction is… See the full description on the dataset page: https://huggingface.co/datasets/kaupane/human-recaption.cc12m-wds-recaption
CC12M with Enhanced Captions
This dataset contains 1.3 million image-text pairs from the CC12M dataset with model-generated captions.
Dataset Details
Total Samples: 1,306,239
Source: pixparse/cc12m-wds
Captioning Model: Qwen/Qwen3-VL-8B-Instruct
Format: Parquet
Filtering Criteria
Samples were filtered based on the following quality metrics:
Aesthetic Score: >= 5.5 (using LAION aesthetic classifier)
Resolution: >= 512 pixels (width or height)
Aspect Ratio: <=… See the full description on the dataset page: https://huggingface.co/datasets/undefined443/cc12m-wds-recaption.LAION-Art-recaption
LAION-Art Recaption
Recaptioned subset of the LAION-Art dataset using Qwen2.5-VL-7B-Instruct.
Dataset Description
This dataset contains detailed recaptions for LAION-Art images generated by a vision-language model. Only successfully recaptioned samples are included.
Total samples: 1,410,704
Columns
Column
Type
Description
url
string
Original image URL
key
string
Unique identifier for each sample
width
int32
Image width
height
int32
Image height… See the full description on the dataset page: https://huggingface.co/datasets/undefined443/LAION-Art-recaption.zangei-dit-stage-1-250k-recaptionedzangei-dit-stage-1-250k-recaptioned-qwen-3-0.6b-embeddings
