datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cc12m-a_woman
Description
This dataset is a convenience subset of https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned/
I did a quick grep for "A woman", and then HAND-CURATED the results.
That means I threw out anything with watermarks, site branding, or pretty much anything else I deemed would
get in the way of ML training.
I also only chose images that had clear, sharp camera focus on the main subject. So these are high-quality images.
At present, I have only done a few thousand.
I… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-a_woman.laion2b-23ish-woman-solo
Overview
All images have a woman in them, solo, at APPROXIMATELY 2:3 aspect ratio.
These images are HUMAN CURATED. I have personally gone through every one at least once.
Additionally, there are no visible watermarks, the quality and focus are good, and it should not be confusing for AI training
There should be a little over 15k images here.
Note that there is a wide variety of body sizes, from size 0, to perhaps size 18
There are also THREE choices of captions: the really bad "alt… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-23ish-woman-solo.pexels-woman-croppable
pexels-woman-croppable
An extract from our larger "pexels 130k images" set. Around 6500 images.
Useful for training text-to-image models.
But we already have a bunch of subsets, why another one?
This is 6k images that, while not originally square cropped, are
HAND-SELECTED to be square-crop clean.
The provided crawl.sh util script will handle automatically
downloading and cropping them from pexels.com
Why Square-Crop
When training a model, you must have all images… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/pexels-woman-croppable.
