datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cc12m-recaptionedcc12m-wds-coco-recaptioned
CC12M WebDataset with COCO-style Recaptions
A large-scale image-text dataset containing 3 million images from Conceptual Captions 12M (CC12M) with COCO-style factual descriptions generated using NVIDIA Nemotron Nano 12B v2 VL.
Dataset Overview
Base Dataset: pixparse/cc12m-wds - Conceptual Captions 12M (CC12M)
Images: 3,000,000+ high-quality internet images
Recaption Model: NVIDIA Nemotron Nano 12B v2 VL
Recaption Style: COCO-style factual descriptions (20 words average)… See the full description on the dataset page: https://huggingface.co/datasets/undefined443/cc12m-wds-coco-recaptioned.Recap-Datacomp-1B_tars_part7Final size: 7,236,721, samples per tar: 10000
cc3m-recap-wdsRecap-Datacomp-1B_tars_part9Final size: 7,240,328, samples per tar: 10000
Recap-Datacomp-1B_tars_part8Final size: 7,250,604, samples per tar: 10000
Recap-Datacomp-1B_tars_part13cc12m_recapRecap-Datacomp-1B_tars_part14cc12m_and_imagenet21k_recap_wdscc12_imagenet21k_recap_hq_bucketed
cc12_imagenet21k_recap_hq_bucketed
Title: cc12_imagenet21k_recap_hq_bucketed
Description: This ~18M rows dataset is a re upload of https://huggingface.co/datasets/gmongaras/CC12M_and_Imagenet21K_Recap_Highqual where the images have
been pre bucketed into SDXL style aspect ratio buckets for target training at ~512^2 and ~256^2 pixels, and where about 7M rows were recaptioned with either Gemini or Ministral.
To avoid re encoding the images they have been left untouched so cropping… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/cc12_imagenet21k_recap_hq_bucketed.recap_data_tarrecap-datacomp-384-1Mcc12m_imagenet21k_recap_256_20m_wds
