datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cc12m-recaptionedcc12m-wds-coco-recaptioned
CC12M WebDataset with COCO-style Recaptions
A large-scale image-text dataset containing 3 million images from Conceptual Captions 12M (CC12M) with COCO-style factual descriptions generated using NVIDIA Nemotron Nano 12B v2 VL.
Dataset Overview
Base Dataset: pixparse/cc12m-wds - Conceptual Captions 12M (CC12M)
Images: 3,000,000+ high-quality internet images
Recaption Model: NVIDIA Nemotron Nano 12B v2 VL
Recaption Style: COCO-style factual descriptions (20 words average)… See the full description on the dataset page: https://huggingface.co/datasets/undefined443/cc12m-wds-coco-recaptioned.Recap-Datacomp-1B_tars_part7Final size: 7,236,721, samples per tar: 10000
cc3m-recap-wdsRecap-Datacomp-1B_tars_part9Final size: 7,240,328, samples per tar: 10000
Recap-Datacomp-1B_tars_part8Final size: 7,250,604, samples per tar: 10000
Recap-Datacomp-1B_tars_part13Recap-Datacomp-1B_tars_part14CyclePrefDB-I2T-Reconstructions
Image Reconstructions for CyclePrefDB-I2T
Project page | Paper | Code
This dataset contains reconstruction images used to determine cycle consistency preferences for CyclePrefDB-I2T. You can find the corresponding file paths in the CyclePrefDB-I2T dataset here. Reconstructions are created using Stable Diffusion 3 Medium.
Preparing the reconstructions
You can download the test and validation split .tar files and extract them directly.Use this script to extract the… See the full description on the dataset page: https://huggingface.co/datasets/carolineec/CyclePrefDB-I2T-Reconstructions.ccpd-ocr-recognitioncc12m_and_imagenet21k_recap_wdscc12_imagenet21k_recap_hq_bucketed
cc12_imagenet21k_recap_hq_bucketed
Title: cc12_imagenet21k_recap_hq_bucketed
Description: This ~18M rows dataset is a re upload of https://huggingface.co/datasets/gmongaras/CC12M_and_Imagenet21K_Recap_Highqual where the images have
been pre bucketed into SDXL style aspect ratio buckets for target training at ~512^2 and ~256^2 pixels, and where about 7M rows were recaptioned with either Gemini or Ministral.
To avoid re encoding the images they have been left untouched so cropping… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/cc12_imagenet21k_recap_hq_bucketed.recap_data_tarcc12m_imagenet21k_recap_256_20m_wdsrecap-datacomp-384-1MOKReddit-Visionary
Dataset Summary
OKReddit Visionary is a collection of 50 GiB (~74K pairs) of image Question & Answers. This dataset has been prepared for research or archival purposes.
Curated by: KaraKaraWitch
Funded by: Recursal.ai
Shared by: KaraKaraWitch
Special Thanks: harrison (Suggestion)
Language(s) (NLP): Mainly English.
License: Refer to Licensing Information for data license.
Dataset Sources
Source Data: Academic Torrents by (stuck_in_the_matrix, Watchful1… See the full description on the dataset page: https://huggingface.co/datasets/recursal/OKReddit-Visionary.cued-recall-imagenet
cued-recall-imagenet
Train/val/test split of ImageNet used by the cued-recall MemoryVLM experiments.
Contents
imagenet/class_XXXX.tar — per-class JPG bundles (235 classes). Extract each
with tar xf class_XXXX.tar to get a class_XXXX/ directory of img_*.jpg.
splits.json — authoritative train/val/test assignment. References paths
relative to data/, e.g. data/imagenet/class_0000/img_00000000.jpg.
memory_datasets.tar — Brady2008/2013 stimulus sets + MST (used for… See the full description on the dataset page: https://huggingface.co/datasets/maxbennett/cued-recall-imagenet.recon_feature
