datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cc12m-structured-captions
CC12M multi-caption Webshart indexes
Metadata-only indexes for the 10,994,853 available images in
laion/conceptual-captions-12m-webdataset, across 1,100 tar shards. The images
remain in that source dataset; this repository does not duplicate or rewrite
the image archives. Original member offsets, lengths, sidecar references, and
image geometry are preserved.
These per-shard Webshart indexes are based on
webshart/conceptual-captions-12m-webdataset-metadata.
The 1,100 JSON files… See the full description on the dataset page: https://huggingface.co/datasets/webshart/cc12m-structured-captions.cottonweed-structured-captions
