CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ooutlierr /cc12m-recaptionedimage10M<n<100M8 likes2.3k downloads2y agoHugging Face02undefined443 /cc12m-wds-coco-recaptioned CC12M WebDataset with COCO-style Recaptions A large-scale image-text dataset containing 3 million images from Conceptual Captions 12M (CC12M) with COCO-style factual descriptions generated using NVIDIA Nemotron Nano 12B v2 VL. Dataset Overview Base Dataset: pixparse/cc12m-wds - Conceptual Captions 12M (CC12M) Images: 3,000,000+ high-quality internet images Recaption Model: NVIDIA Nemotron Nano 12B v2 VL Recaption Style: COCO-style factual descriptions (20 words average)… See the full description on the dataset page: https://huggingface.co/datasets/undefined443/cc12m-wds-coco-recaptioned.image1M<n<10M1 likes1.4k downloads6mo agoHugging Face03wjn922 /Recap-Datacomp-1B_tars_part7Final size: 7,236,721, samples per tar: 10000 image1M<n<10M0 likes335 downloads11mo agoHugging Face04Leonardo6 /cc3m-recap-wdsimage1M<n<10M0 likes327 downloads1y agoHugging Face05wjn922 /Recap-Datacomp-1B_tars_part9Final size: 7,240,328, samples per tar: 10000 image1M<n<10M0 likes205 downloads11mo agoHugging Face06wjn922 /Recap-Datacomp-1B_tars_part8Final size: 7,250,604, samples per tar: 10000 image1M<n<10M0 likes180 downloads11mo agoHugging Face07wjn922 /Recap-Datacomp-1B_tars_part13image1M<n<10M0 likes176 downloads11mo agoHugging Face08wjn922 /Recap-Datacomp-1B_tars_part14image1M<n<10M0 likes158 downloads11mo agoHugging Face09carolineec /CyclePrefDB-I2T-Reconstructions Image Reconstructions for CyclePrefDB-I2T Project page | Paper | Code This dataset contains reconstruction images used to determine cycle consistency preferences for CyclePrefDB-I2T. You can find the corresponding file paths in the CyclePrefDB-I2T dataset here. Reconstructions are created using Stable Diffusion 3 Medium. Preparing the reconstructions You can download the test and validation split .tar files and extract them directly.Use this script to extract the… See the full description on the dataset page: https://huggingface.co/datasets/carolineec/CyclePrefDB-I2T-Reconstructions.imageimage-to-text10K<n<100K0 likes108 downloads1y agoHugging Face10zenitsu09 /ccpd-ocr-recognitionimage10K<n<100K1 likes107 downloads1y agoHugging Face11Leonardo6 /cc12m_and_imagenet21k_recap_wdsimage1M<n<10M0 likes62 downloads1y agoHugging Face12data-archetype /cc12_imagenet21k_recap_hq_bucketed cc12_imagenet21k_recap_hq_bucketed Title: cc12_imagenet21k_recap_hq_bucketed Description: This ~18M rows dataset is a re upload of https://huggingface.co/datasets/gmongaras/CC12M_and_Imagenet21K_Recap_Highqual where the images have been pre bucketed into SDXL style aspect ratio buckets for target training at ~512^2 and ~256^2 pixels, and where about 7M rows were recaptioned with either Gemini or Ministral. To avoid re encoding the images they have been left untouched so cropping… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/cc12_imagenet21k_recap_hq_bucketed.imagetext-to-image10M<n<100M0 likes59 downloads9mo agoHugging Face13wudifanfandou /recap_data_tarimage10M<n<100M0 likes21 downloads1y agoHugging Face14Leonardo6 /cc12m_imagenet21k_recap_256_20m_wdsimage1M<n<10M0 likes19 downloads1y agoHugging Face15wooj1nBot /recap-datacomp-384-1Mimage100K<n<1M0 likes18 downloads1y agoHugging Face16recursal /OKReddit-Visionary Dataset Summary OKReddit Visionary is a collection of 50 GiB (~74K pairs) of image Question & Answers. This dataset has been prepared for research or archival purposes. Curated by: KaraKaraWitch Funded by: Recursal.ai Shared by: KaraKaraWitch Special Thanks: harrison (Suggestion) Language(s) (NLP): Mainly English. License: Refer to Licensing Information for data license. Dataset Sources Source Data: Academic Torrents by (stuck_in_the_matrix, Watchful1… See the full description on the dataset page: https://huggingface.co/datasets/recursal/OKReddit-Visionary.imagequestion-answering100K<n<1M1 likes11 downloads2y agoHugging Face17maxbennett /cued-recall-imagenet cued-recall-imagenet Train/val/test split of ImageNet used by the cued-recall MemoryVLM experiments. Contents imagenet/class_XXXX.tar — per-class JPG bundles (235 classes). Extract each with tar xf class_XXXX.tar to get a class_XXXX/ directory of img_*.jpg. splits.json — authoritative train/val/test assignment. References paths relative to data/, e.g. data/imagenet/class_0000/img_00000000.jpg. memory_datasets.tar — Brady2008/2013 stimulus sets + MST (used for… See the full description on the dataset page: https://huggingface.co/datasets/maxbennett/cued-recall-imagenet.image100K<n<1M0 likes10 downloads5mo agoHugging Face18l3ss /recon_featureimage10K<n<100K0 likes3 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.