datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cc12m-cleaned
CC12m-cleaned
This dataset builds on two others: The Conceptual Captions 12million dataset, which lead to the LLaVa captioned subset done by
CaptionEmporium
(The latter is the same set, but swaps out the (Conceptual Captions 12million) often-useless alt-text captioning for decent ones_
I have then used the llava captions as a base, and used the detailed descrptions to filter out
images with things like watermarks, artist signatures, etc.
I have also manually thrown out all… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned.conceptual-captions-cc12m-llavanext
Dataset Card for conceptual-captions-cc12m-llavanext
Dataset Summary
This is a data of 21,930,344 synthetic captions for 10,965,172 images from conceptual_12m. In the interest of reproducibility, an archive found here on Huggingface was used (cc12m-wds). The captions were produced using llama3-llava-next-8b inferenced in float16, followed by cleanup and shortening with Meta-Llama-3-8B.
Languages
The captions are in English.
Data Instances
An… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/conceptual-captions-cc12m-llavanext.CC_12M_Indonesiacc12m-a_woman
Description
This dataset is a convenience subset of https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned/
I did a quick grep for "A woman", and then HAND-CURATED the results.
That means I threw out anything with watermarks, site branding, or pretty much anything else I deemed would
get in the way of ML training.
I also only chose images that had clear, sharp camera focus on the main subject. So these are high-quality images.
At present, I have only done a few thousand.
I… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-a_woman.cc12m-4mp-realistic
Overview
This is a hand-selected subset of our larger attempts to filter the well known CC12M dataset.
This one focuses on large (4 megapixels) images that are real world, high quality images, and the captioning
specifically matches either "A man" or "A woman".
Note that I did not have the diskspace/time to go through the ENTIRE set. It was perhaps only from the first 2 million of our
CC12M-cleaned subset.
If an effort were made to go through the entire 4mp image set, there might be… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-4mp-realistic.cc12m-2mp-realistic
Overview
A subset of the "CC12m" dataset. Size varying between 2mp <= x < 4mp. Anything actually 4mp or larger will be in one of our "4mp" datasets.
This dataset is created for if you basically need more images and dont mind a little less quality.
Quality
I have filtered out as many watermarks, etc. as possible using AI models. I have also thrown out stupid black-and-white photos, because they poison normal image prompting.
This is NOT HAND CURATED, unlike some… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-2mp-realistic.cc12m-1mp_plus-realistic
cc12m-1mp_plus-realistic
A filtering down of the full CC12M dataset, to have the following characteristics:
At least 1024x1024 pixels in size
"Realistic". No paintings, digital art, monochrome, or surreal stuff. Also discard multi-image as much as possible
Ideally, no signed or watermarked images. (but there will certainly be some left)
Captions
The caption types available are a bit different from some of our other ones. Currently available are:
caption_llava… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-1mp_plus-realistic.cc12m-4mp
Parent dataset
https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned
Contents
I had uploaded a prior version of this "4mp" set, that was just partial. Now I have grabbed all of them available as of 2024 December.
This latest upload references ALL files in the "CC12M" dataset, that are at least 4 megapixel in size,
and are not one of the known "stock images" copyrighted sites. (and were available this month)
I have done some additional filtering out of any… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-4mp.cc12m-small-squarish-simple
Contents of dataset
This is a "general purposes dataset", of images that are 512px to 1024px in size, if I recall correctly
(in contrast to the "2mp" and "4mp" datasets)
Also, they are "squarish", which means that, even if they are not precisely square, the image looks fine if you
hard-crop it to force square.
Additionally, images have been hand-culled to throw out anything I considered bad for AI training.
Additionally, they were AI-culled to have a "simple photographic… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-small-squarish-simple.cc12m-2mp-squareish
Overview
A subset of the "CC12m" dataset. Around 37k images, varying between 2mp <= x < 4mp.
Anything actually 4mp or larger will be in one of our "4mp" datasets. This dataset is created for
if you basically need more images and dont mind a little less quality.
Why "Squareish"?
I noticed that SD1.5 really, really likes square images to train on.
These are mostly NOT EXACTLY SQUARE. However, most training programs should have an auto-crop capability.
This dataset is of… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-2mp-squareish.cc12m-xlsdA quick upload of my current experimental dataset for XLSD.
Approximately 350k images, ranging between 2mp and 4mp in size. Downloaded image disk space is somewhere under 400G
An attempt has been made via AI to remove black and white images, non-real images, and watermarks, etc.
So this is basically a "realistic photo" image set, with aspect ratios ranging from 3:2 to 2:3.
We do not include 16:9 here, since I figured that would be a waste of information for a general purpose SD unet.
This… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-xlsd.cc12m-1mp_plus-realistic
cc12m-1mp_plus-realistic
A filtering down of the full CC12M dataset, to have the following characteristics:
At least 1024x1024 pixels in size
"Realistic". No paintings, digital art, monochrome, or surreal stuff. Also discard multi-image as much as possible
Ideally, no signed or watermarked images. (but there will certainly be some left)
Captions
The caption types available are a bit different from some of our other ones. Currently available are:
caption_llava… See the full description on the dataset page: https://huggingface.co/datasets/QinboZhang/cc12m-1mp_plus-realistic.
