datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
laion2b-23ish-woman-solo
Overview
All images have a woman in them, solo, at APPROXIMATELY 2:3 aspect ratio.
These images are HUMAN CURATED. I have personally gone through every one at least once.
Additionally, there are no visible watermarks, the quality and focus are good, and it should not be confusing for AI training
There should be a little over 15k images here.
Note that there is a wide variety of body sizes, from size 0, to perhaps size 18
There are also THREE choices of captions: the really bad "alt… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-23ish-woman-solo.laion2b-en-aesthetic-square-human
Overview
This dataset is a HAND-CURATED version of our laion2b-en-aesthetic-square-cleaned dataset.
It has at least "a man" or "a woman" in it.
Additionally, there are no visible watermarks, the quality and focus are good, and it should not be confusing for AI training
There should be a little over 8k images here.
Details
It consists of an initial extraction of all images that had "a man" or "a woman" in the moondream caption.
I then filtered out all "statue" or… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-en-aesthetic-square-human.laion2b-en-aesthetic-square-cleaned
Overview
A subset of our opendiffusionai/laion2b-en-aesthetic-square, which is itself a subset of the widely known "Laion2b-en-aesthetic" dataset.
However, the original had only the website alt-tags for captions.
I have added decent AI captioning, via the "moondream2" model.
Additionally, I have stripped out a bunch of watermarked junk, and weeded out 20k duplicate images.
When I use it for training, I will be trimming out additional things like non-realistic images, painting… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-en-aesthetic-square-cleaned.laion2b-mixed-1024px-human
Overview
This dataset is a selective merge of some other of our datasets. Mainly, I pulled human-centric, real-world photos from the following datasets:
opendiffusionai/laion2b-45ish-1120px
opendiffusionai/laion2b-squareish-1024px
opendiffusionai/laion2b-23ish-1216px
As such, it is "mixed" aspect ratio.
The very smallest height ones are from our "squarish" set, so are at least 1024px tall.
However, the other ones with longer rations, have an appropriately longer minimum… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-mixed-1024px-human.laion2b-squareish-1536px
Overview
This dataset is a subset of laion2b-en-aesthetic dataset. Its purpose is for general-case (photographic) model training.
Its aspect ratios are (1 <= ratio < 5/4 ) so that images can be auto-cropped to square safely.
Additionally, all images are at least 1536 pixels tall.
This is because I wanted a dataset that would be very high quality to use for models that are 768x768
There should be close to 80k 8k images here.
Note that you can CHOOSE between "moondream"… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-squareish-1536px.laion2b-squareish-1024px
Overview
This dataset is a subset of laion2b-en-aesthetic dataset. Its purpose is for general-case (photographic) model training.
Its aspect ratios are (1 <= ratio < 5/4 ) so that images can be auto-cropped to square safely. Additionally, all images are at least 1024 pixels tall.
There should be close to 250k images here. Total on-disk size is approximately 118G
We have a very similar dataset that is 1536px in size. However, that is only 80k in size, so if you need waaay more… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-squareish-1024px.laion2b-45ish-1120px
Overview
This is a subset of https://huggingface.co/datasets/laion/laion2B-en-aesthetic, selected for aspect ratio, and with better captioning.
It is a GENERAL CASE dataset. Percentages of images with humans in it, is approximately only 30%
I also filtered out all non-realistic images. This is intended to be a "real world" dataset.
Approximate image count in this dataset is around 80k.
On-disk size is around 45G.
This is NOT individually human filtered. Batch culling only.… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-45ish-1120px.laion2b-squareish
General purpose dataset (realistic)
This is a merging of our laion2b-squareish-1024px and laion2b-squareish-1536px datasets,
WITH some additional filtering.
It has a little de-duplication, and some extra watermark removal.
It is MOSTLY realistic.
I'm sharing this, because I am actively using it in my SD1.5 training experiments.
I needed specifically square(ish) rather than our other nice, but mixed aspect-ratio datasets.
Plus, I needed a LARGE one.
So, here it is!… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-squareish.laion2b-en-aesthetic-square
Contents
This is a pretty raw filter of https://huggingface.co/datasets/laion/laion2B-en-aesthetic
I just filtered for "is image perfectly square, AND is image at least 1024x1024 pixels"
Approximate image count is a bit over 300k
Updated 2025/01/24
I just found out there are a bunch of watermarked sites in here.
So much for aesthetically chosen :(
So I filtered out a bunch of the "stock image" sites, just by looking at url strings.
Update 2025/01/31… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-en-aesthetic-square.laion2b-34ish-1152pxDont use this yet. Work in progress.
In theory you could, but its messy data.
I uploaded it here so I could more easily transfer it off my home machine to a research machine, where I will then filter it for
watermarks and give it better captions.
Should take maybe a week.
laion2b-23ish-1216px
Overview
This is a subset of https://huggingface.co/datasets/laion/laion2B-en-aesthetic,
selected for aspect ratio, and with better captioning.
Approximate image count is around 250k.
23ish, 1216px
I picked out the images that are portrait aspect ratio of 2:3, or a little wider (Because images that are a little too wide, can be safely cropped narrower)
I also picked a minimum height of 1216 pixels, because that is what 1024x1024 pixelcount converted to 2:3 looks like.… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-23ish-1216px.laion2b-43ish-1000px
Overview
This is a subset of https://huggingface.co/datasets/laion/laion2B-en-aesthetic, selected for aspect ratio.
It is a GENERAL CASE dataset.
I attempted to filter out non-realistic images. This is intended to be a "real world" dataset.
Approximate image count in this dataset is around 170k. On-disk size is around 185G.
This is NOT individually human filtered. Almost no culling.
That being said, the images are still relatively high quality!
Approximately 4:3ish aspect ratio… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-43ish-1000px.
