CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ooutlierr /cc12m-recaptionedimage10M<n<100M8 likes2.1k downloads2y agoHugging Face02undefined443 /cc12m-wds-coco-recaptioned CC12M WebDataset with COCO-style Recaptions A large-scale image-text dataset containing 3 million images from Conceptual Captions 12M (CC12M) with COCO-style factual descriptions generated using NVIDIA Nemotron Nano 12B v2 VL. Dataset Overview Base Dataset: pixparse/cc12m-wds - Conceptual Captions 12M (CC12M) Images: 3,000,000+ high-quality internet images Recaption Model: NVIDIA Nemotron Nano 12B v2 VL Recaption Style: COCO-style factual descriptions (20 words average)… See the full description on the dataset page: https://huggingface.co/datasets/undefined443/cc12m-wds-coco-recaptioned.image1M<n<10M1 likes1.6k downloads5mo agoHugging Face03wjn922 /Recap-Datacomp-1B_tars_part7Final size: 7,236,721, samples per tar: 10000 image1M<n<10M0 likes337 downloads11mo agoHugging Face04Leonardo6 /cc3m-recap-wdsimage1M<n<10M0 likes327 downloads1y agoHugging Face05wjn922 /Recap-Datacomp-1B_tars_part9Final size: 7,240,328, samples per tar: 10000 image1M<n<10M0 likes205 downloads11mo agoHugging Face06wjn922 /Recap-Datacomp-1B_tars_part8Final size: 7,250,604, samples per tar: 10000 image1M<n<10M0 likes180 downloads11mo agoHugging Face07wjn922 /Recap-Datacomp-1B_tars_part13image1M<n<10M0 likes175 downloads10mo agoHugging Face08wusize /cc12m_recaptext1M<n<10M0 likes168 downloads1y agoHugging Face09wjn922 /Recap-Datacomp-1B_tars_part14image1M<n<10M0 likes158 downloads10mo agoHugging Face10Leonardo6 /cc12m_and_imagenet21k_recap_wdsimage1M<n<10M0 likes63 downloads1y agoHugging Face11data-archetype /cc12_imagenet21k_recap_hq_bucketed cc12_imagenet21k_recap_hq_bucketed Title: cc12_imagenet21k_recap_hq_bucketed Description: This ~18M rows dataset is a re upload of https://huggingface.co/datasets/gmongaras/CC12M_and_Imagenet21K_Recap_Highqual where the images have been pre bucketed into SDXL style aspect ratio buckets for target training at ~512^2 and ~256^2 pixels, and where about 7M rows were recaptioned with either Gemini or Ministral. To avoid re encoding the images they have been left untouched so cropping… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/cc12_imagenet21k_recap_hq_bucketed.imagetext-to-image10M<n<100M0 likes57 downloads9mo agoHugging Face12wudifanfandou /recap_data_tarimage10M<n<100M0 likes21 downloads1y agoHugging Face13wooj1nBot /recap-datacomp-384-1Mimage100K<n<1M0 likes18 downloads1y agoHugging Face14Leonardo6 /cc12m_imagenet21k_recap_256_20m_wdsimage1M<n<10M0 likes16 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.