datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cc12m-cleaned
CC12m-cleaned
This dataset builds on two others: The Conceptual Captions 12million dataset, which lead to the LLaVa captioned subset done by
CaptionEmporium
(The latter is the same set, but swaps out the (Conceptual Captions 12million) often-useless alt-text captioning for decent ones_
I have then used the llava captions as a base, and used the detailed descrptions to filter out
images with things like watermarks, artist signatures, etc.
I have also manually thrown out all… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned.cc12m-a_woman
Description
This dataset is a convenience subset of https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned/
I did a quick grep for "A woman", and then HAND-CURATED the results.
That means I threw out anything with watermarks, site branding, or pretty much anything else I deemed would
get in the way of ML training.
I also only chose images that had clear, sharp camera focus on the main subject. So these are high-quality images.
At present, I have only done a few thousand.
I… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-a_woman.laion2b-23ish-woman-solo
Overview
All images have a woman in them, solo, at APPROXIMATELY 2:3 aspect ratio.
These images are HUMAN CURATED. I have personally gone through every one at least once.
Additionally, there are no visible watermarks, the quality and focus are good, and it should not be confusing for AI training
There should be a little over 15k images here.
Note that there is a wide variety of body sizes, from size 0, to perhaps size 18
There are also THREE choices of captions: the really bad "alt… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-23ish-woman-solo.laion2b-en-aesthetic-square-human
Overview
This dataset is a HAND-CURATED version of our laion2b-en-aesthetic-square-cleaned dataset.
It has at least "a man" or "a woman" in it.
Additionally, there are no visible watermarks, the quality and focus are good, and it should not be confusing for AI training
There should be a little over 8k images here.
Details
It consists of an initial extraction of all images that had "a man" or "a woman" in the moondream caption.
I then filtered out all "statue" or… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-en-aesthetic-square-human.cc12m-4mp-realistic
Overview
This is a hand-selected subset of our larger attempts to filter the well known CC12M dataset.
This one focuses on large (4 megapixels) images that are real world, high quality images, and the captioning
specifically matches either "A man" or "A woman".
Note that I did not have the diskspace/time to go through the ENTIRE set. It was perhaps only from the first 2 million of our
CC12M-cleaned subset.
If an effort were made to go through the entire 4mp image set, there might be… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-4mp-realistic.cc12m-2mp-realistic
Overview
A subset of the "CC12m" dataset. Size varying between 2mp <= x < 4mp. Anything actually 4mp or larger will be in one of our "4mp" datasets.
This dataset is created for if you basically need more images and dont mind a little less quality.
Quality
I have filtered out as many watermarks, etc. as possible using AI models. I have also thrown out stupid black-and-white photos, because they poison normal image prompting.
This is NOT HAND CURATED, unlike some… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-2mp-realistic.laion2b-en-aesthetic-square-cleaned
Overview
A subset of our opendiffusionai/laion2b-en-aesthetic-square, which is itself a subset of the widely known "Laion2b-en-aesthetic" dataset.
However, the original had only the website alt-tags for captions.
I have added decent AI captioning, via the "moondream2" model.
Additionally, I have stripped out a bunch of watermarked junk, and weeded out 20k duplicate images.
When I use it for training, I will be trimming out additional things like non-realistic images, painting… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-en-aesthetic-square-cleaned.cc12m-1mp_plus-realistic
cc12m-1mp_plus-realistic
A filtering down of the full CC12M dataset, to have the following characteristics:
At least 1024x1024 pixels in size
"Realistic". No paintings, digital art, monochrome, or surreal stuff. Also discard multi-image as much as possible
Ideally, no signed or watermarked images. (but there will certainly be some left)
Captions
The caption types available are a bit different from some of our other ones. Currently available are:
caption_llava… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-1mp_plus-realistic.pexels-woman-croppable
pexels-woman-croppable
An extract from our larger "pexels 130k images" set. Around 6500 images.
Useful for training text-to-image models.
But we already have a bunch of subsets, why another one?
This is 6k images that, while not originally square cropped, are
HAND-SELECTED to be square-crop clean.
The provided crawl.sh util script will handle automatically
downloading and cropping them from pexels.com
Why Square-Crop
When training a model, you must have all images… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/pexels-woman-croppable.cc12m-4mp
Parent dataset
https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned
Contents
I had uploaded a prior version of this "4mp" set, that was just partial. Now I have grabbed all of them available as of 2024 December.
This latest upload references ALL files in the "CC12M" dataset, that are at least 4 megapixel in size,
and are not one of the known "stock images" copyrighted sites. (and were available this month)
I have done some additional filtering out of any… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-4mp.laion2b-mixed-1024px-human
Overview
This dataset is a selective merge of some other of our datasets. Mainly, I pulled human-centric, real-world photos from the following datasets:
opendiffusionai/laion2b-45ish-1120px
opendiffusionai/laion2b-squareish-1024px
opendiffusionai/laion2b-23ish-1216px
As such, it is "mixed" aspect ratio.
The very smallest height ones are from our "squarish" set, so are at least 1024px tall.
However, the other ones with longer rations, have an appropriately longer minimum… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-mixed-1024px-human.laion2b-squareish-1536px
Overview
This dataset is a subset of laion2b-en-aesthetic dataset. Its purpose is for general-case (photographic) model training.
Its aspect ratios are (1 <= ratio < 5/4 ) so that images can be auto-cropped to square safely.
Additionally, all images are at least 1536 pixels tall.
This is because I wanted a dataset that would be very high quality to use for models that are 768x768
There should be close to 80k 8k images here.
Note that you can CHOOSE between "moondream"… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-squareish-1536px.laion2b-squareish-1024px
Overview
This dataset is a subset of laion2b-en-aesthetic dataset. Its purpose is for general-case (photographic) model training.
Its aspect ratios are (1 <= ratio < 5/4 ) so that images can be auto-cropped to square safely. Additionally, all images are at least 1024 pixels tall.
There should be close to 250k images here. Total on-disk size is approximately 118G
We have a very similar dataset that is 1536px in size. However, that is only 80k in size, so if you need waaay more… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-squareish-1024px.laion2b-45ish-1120px
Overview
This is a subset of https://huggingface.co/datasets/laion/laion2B-en-aesthetic, selected for aspect ratio, and with better captioning.
It is a GENERAL CASE dataset. Percentages of images with humans in it, is approximately only 30%
I also filtered out all non-realistic images. This is intended to be a "real world" dataset.
Approximate image count in this dataset is around 80k.
On-disk size is around 45G.
This is NOT individually human filtered. Batch culling only.… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-45ish-1120px.laion2b-squareish
General purpose dataset (realistic)
This is a merging of our laion2b-squareish-1024px and laion2b-squareish-1536px datasets,
WITH some additional filtering.
It has a little de-duplication, and some extra watermark removal.
It is MOSTLY realistic.
I'm sharing this, because I am actively using it in my SD1.5 training experiments.
I needed specifically square(ish) rather than our other nice, but mixed aspect-ratio datasets.
Plus, I needed a LARGE one.
So, here it is!… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-squareish.cc12m-small-squarish-simple
Contents of dataset
This is a "general purposes dataset", of images that are 512px to 1024px in size, if I recall correctly
(in contrast to the "2mp" and "4mp" datasets)
Also, they are "squarish", which means that, even if they are not precisely square, the image looks fine if you
hard-crop it to force square.
Additionally, images have been hand-culled to throw out anything I considered bad for AI training.
Additionally, they were AI-culled to have a "simple photographic… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-small-squarish-simple.laion2b-en-aesthetic-square
Contents
This is a pretty raw filter of https://huggingface.co/datasets/laion/laion2B-en-aesthetic
I just filtered for "is image perfectly square, AND is image at least 1024x1024 pixels"
Approximate image count is a bit over 300k
Updated 2025/01/24
I just found out there are a bunch of watermarked sites in here.
So much for aesthetically chosen :(
So I filtered out a bunch of the "stock image" sites, just by looking at url strings.
Update 2025/01/31… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-en-aesthetic-square.cc12m-2mp-squareish
Overview
A subset of the "CC12m" dataset. Around 37k images, varying between 2mp <= x < 4mp.
Anything actually 4mp or larger will be in one of our "4mp" datasets. This dataset is created for
if you basically need more images and dont mind a little less quality.
Why "Squareish"?
I noticed that SD1.5 really, really likes square images to train on.
These are mostly NOT EXACTLY SQUARE. However, most training programs should have an auto-crop capability.
This dataset is of… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-2mp-squareish.laion2b-34ish-1152pxDont use this yet. Work in progress.
In theory you could, but its messy data.
I uploaded it here so I could more easily transfer it off my home machine to a research machine, where I will then filter it for
watermarks and give it better captions.
Should take maybe a week.
laion2b-23ish-1216px
Overview
This is a subset of https://huggingface.co/datasets/laion/laion2B-en-aesthetic,
selected for aspect ratio, and with better captioning.
Approximate image count is around 250k.
23ish, 1216px
I picked out the images that are portrait aspect ratio of 2:3, or a little wider (Because images that are a little too wide, can be safely cropped narrower)
I also picked a minimum height of 1216 pixels, because that is what 1024x1024 pixelcount converted to 2:3 looks like.… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-23ish-1216px.cc12m-xlsdA quick upload of my current experimental dataset for XLSD.
Approximately 350k images, ranging between 2mp and 4mp in size. Downloaded image disk space is somewhere under 400G
An attempt has been made via AI to remove black and white images, non-real images, and watermarks, etc.
So this is basically a "realistic photo" image set, with aspect ratios ranging from 3:2 to 2:3.
We do not include 16:9 here, since I figured that would be a waste of information for a general purpose SD unet.
This… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-xlsd.laion2b-43ish-1000px
Overview
This is a subset of https://huggingface.co/datasets/laion/laion2B-en-aesthetic, selected for aspect ratio.
It is a GENERAL CASE dataset.
I attempted to filter out non-realistic images. This is intended to be a "real world" dataset.
Approximate image count in this dataset is around 170k. On-disk size is around 185G.
This is NOT individually human filtered. Almost no culling.
That being said, the images are still relatively high quality!
Approximately 4:3ish aspect ratio… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-43ish-1000px.
