CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01opendiffusionai /cc12m-cleaned CC12m-cleaned This dataset builds on two others: The Conceptual Captions 12million dataset, which lead to the LLaVa captioned subset done by CaptionEmporium (The latter is the same set, but swaps out the (Conceptual Captions 12million) often-useless alt-text captioning for decent ones_ I have then used the llava captions as a base, and used the detailed descrptions to filter out images with things like watermarks, artist signatures, etc. I have also manually thrown out all… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned.imagetext-to-image1M<n<10M13 likes269 downloads2y agoHugging Face02opendiffusionai /cc12m-a_woman Description This dataset is a convenience subset of https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned/ I did a quick grep for "A woman", and then HAND-CURATED the results. That means I threw out anything with watermarks, site branding, or pretty much anything else I deemed would get in the way of ML training. I also only chose images that had clear, sharp camera focus on the main subject. So these are high-quality images. At present, I have only done a few thousand. I… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-a_woman.image10K<n<100K27 likes79 downloads2y agoHugging Face03opendiffusionai /laion2b-23ish-woman-solo Overview All images have a woman in them, solo, at APPROXIMATELY 2:3 aspect ratio. These images are HUMAN CURATED. I have personally gone through every one at least once. Additionally, there are no visible watermarks, the quality and focus are good, and it should not be confusing for AI training There should be a little over 15k images here. Note that there is a wide variety of body sizes, from size 0, to perhaps size 18 There are also THREE choices of captions: the really bad "alt… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-23ish-woman-solo.image10K<n<100K19 likes70 downloads1y agoHugging Face04opendiffusionai /laion2b-en-aesthetic-square-human Overview This dataset is a HAND-CURATED version of our laion2b-en-aesthetic-square-cleaned dataset. It has at least "a man" or "a woman" in it. Additionally, there are no visible watermarks, the quality and focus are good, and it should not be confusing for AI training There should be a little over 8k images here. Details It consists of an initial extraction of all images that had "a man" or "a woman" in the moondream caption. I then filtered out all "statue" or… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-en-aesthetic-square-human.image1K<n<10K8 likes66 downloads2y agoHugging Face05opendiffusionai /cc12m-4mp-realistic Overview This is a hand-selected subset of our larger attempts to filter the well known CC12M dataset. This one focuses on large (4 megapixels) images that are real world, high quality images, and the captioning specifically matches either "A man" or "A woman". Note that I did not have the diskspace/time to go through the ENTIRE set. It was perhaps only from the first 2 million of our CC12M-cleaned subset. If an effort were made to go through the entire 4mp image set, there might be… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-4mp-realistic.imagetext-to-image10K<n<100K22 likes53 downloads2y agoHugging Face06opendiffusionai /cc12m-2mp-realistic Overview A subset of the "CC12m" dataset. Size varying between 2mp <= x < 4mp. Anything actually 4mp or larger will be in one of our "4mp" datasets. This dataset is created for if you basically need more images and dont mind a little less quality. Quality I have filtered out as many watermarks, etc. as possible using AI models. I have also thrown out stupid black-and-white photos, because they poison normal image prompting. This is NOT HAND CURATED, unlike some… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-2mp-realistic.image100K<n<1M4 likes42 downloads2y agoHugging Face07opendiffusionai /laion2b-en-aesthetic-square-cleaned Overview A subset of our opendiffusionai/laion2b-en-aesthetic-square, which is itself a subset of the widely known "Laion2b-en-aesthetic" dataset. However, the original had only the website alt-tags for captions. I have added decent AI captioning, via the "moondream2" model. Additionally, I have stripped out a bunch of watermarked junk, and weeded out 20k duplicate images. When I use it for training, I will be trimming out additional things like non-realistic images, painting… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-en-aesthetic-square-cleaned.image100K<n<1M11 likes38 downloads8mo agoHugging Face08opendiffusionai /cc12m-1mp_plus-realistic cc12m-1mp_plus-realistic A filtering down of the full CC12M dataset, to have the following characteristics: At least 1024x1024 pixels in size "Realistic". No paintings, digital art, monochrome, or surreal stuff. Also discard multi-image as much as possible Ideally, no signed or watermarked images. (but there will certainly be some left) Captions The caption types available are a bit different from some of our other ones. Currently available are: caption_llava… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-1mp_plus-realistic.image100K<n<1M4 likes31 downloads10mo agoHugging Face09opendiffusionai /pexels-woman-croppable pexels-woman-croppable An extract from our larger "pexels 130k images" set. Around 6500 images. Useful for training text-to-image models. But we already have a bunch of subsets, why another one? This is 6k images that, while not originally square cropped, are HAND-SELECTED to be square-crop clean. The provided crawl.sh util script will handle automatically downloading and cropping them from pexels.com Why Square-Crop When training a model, you must have all images… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/pexels-woman-croppable.imagetext-to-image1K<n<10K0 likes24 downloads8mo agoHugging Face10opendiffusionai /cc12m-4mp Parent dataset https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned Contents I had uploaded a prior version of this "4mp" set, that was just partial. Now I have grabbed all of them available as of 2024 December. This latest upload references ALL files in the "CC12M" dataset, that are at least 4 megapixel in size, and are not one of the known "stock images" copyrighted sites. (and were available this month) I have done some additional filtering out of any… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-4mp.image100K<n<1M3 likes21 downloads2y agoHugging Face11opendiffusionai /laion2b-mixed-1024px-human Overview This dataset is a selective merge of some other of our datasets. Mainly, I pulled human-centric, real-world photos from the following datasets: opendiffusionai/laion2b-45ish-1120px opendiffusionai/laion2b-squareish-1024px opendiffusionai/laion2b-23ish-1216px As such, it is "mixed" aspect ratio. The very smallest height ones are from our "squarish" set, so are at least 1024px tall. However, the other ones with longer rations, have an appropriately longer minimum… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-mixed-1024px-human.image10K<n<100K3 likes21 downloads2y agoHugging Face12opendiffusionai /laion2b-squareish-1536px Overview This dataset is a subset of laion2b-en-aesthetic dataset. Its purpose is for general-case (photographic) model training. Its aspect ratios are (1 <= ratio < 5/4 ) so that images can be auto-cropped to square safely. Additionally, all images are at least 1536 pixels tall. This is because I wanted a dataset that would be very high quality to use for models that are 768x768 There should be close to 80k 8k images here. Note that you can CHOOSE between "moondream"… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-squareish-1536px.image10K<n<100K3 likes19 downloads2y agoHugging Face13opendiffusionai /laion2b-squareish-1024px Overview This dataset is a subset of laion2b-en-aesthetic dataset. Its purpose is for general-case (photographic) model training. Its aspect ratios are (1 <= ratio < 5/4 ) so that images can be auto-cropped to square safely. Additionally, all images are at least 1024 pixels tall. There should be close to 250k images here. Total on-disk size is approximately 118G We have a very similar dataset that is 1536px in size. However, that is only 80k in size, so if you need waaay more… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-squareish-1024px.image100K<n<1M1 likes19 downloads2y agoHugging Face14opendiffusionai /laion2b-45ish-1120px Overview This is a subset of https://huggingface.co/datasets/laion/laion2B-en-aesthetic, selected for aspect ratio, and with better captioning. It is a GENERAL CASE dataset. Percentages of images with humans in it, is approximately only 30% I also filtered out all non-realistic images. This is intended to be a "real world" dataset. Approximate image count in this dataset is around 80k. On-disk size is around 45G. This is NOT individually human filtered. Batch culling only.… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-45ish-1120px.image10K<n<100K3 likes17 downloads2y agoHugging Face15opendiffusionai /laion2b-squareish General purpose dataset (realistic) This is a merging of our laion2b-squareish-1024px and laion2b-squareish-1536px datasets, WITH some additional filtering. It has a little de-duplication, and some extra watermark removal. It is MOSTLY realistic. I'm sharing this, because I am actively using it in my SD1.5 training experiments. I needed specifically square(ish) rather than our other nice, but mixed aspect-ratio datasets. Plus, I needed a LARGE one. So, here it is!… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-squareish.image100K<n<1M0 likes13 downloads1y agoHugging Face16opendiffusionai /cc12m-small-squarish-simple Contents of dataset This is a "general purposes dataset", of images that are 512px to 1024px in size, if I recall correctly (in contrast to the "2mp" and "4mp" datasets) Also, they are "squarish", which means that, even if they are not precisely square, the image looks fine if you hard-crop it to force square. Additionally, images have been hand-culled to throw out anything I considered bad for AI training. Additionally, they were AI-culled to have a "simple photographic… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-small-squarish-simple.image10K<n<100K0 likes9 downloads1y agoHugging Face17opendiffusionai /laion2b-en-aesthetic-square Contents This is a pretty raw filter of https://huggingface.co/datasets/laion/laion2B-en-aesthetic I just filtered for "is image perfectly square, AND is image at least 1024x1024 pixels" Approximate image count is a bit over 300k Updated 2025/01/24 I just found out there are a bunch of watermarked sites in here. So much for aesthetically chosen :( So I filtered out a bunch of the "stock image" sites, just by looking at url strings. Update 2025/01/31… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-en-aesthetic-square.image100K<n<1M1 likes8 downloads2y agoHugging Face18opendiffusionai /cc12m-2mp-squareish Overview A subset of the "CC12m" dataset. Around 37k images, varying between 2mp <= x < 4mp. Anything actually 4mp or larger will be in one of our "4mp" datasets. This dataset is created for if you basically need more images and dont mind a little less quality. Why "Squareish"? I noticed that SD1.5 really, really likes square images to train on. These are mostly NOT EXACTLY SQUARE. However, most training programs should have an auto-crop capability. This dataset is of… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-2mp-squareish.image10K<n<100K0 likes6 downloads2y agoHugging Face19opendiffusionai /laion2b-34ish-1152pxDont use this yet. Work in progress. In theory you could, but its messy data. I uploaded it here so I could more easily transfer it off my home machine to a research machine, where I will then filter it for watermarks and give it better captions. Should take maybe a week. image100K<n<1M0 likes6 downloads10mo agoHugging Face20opendiffusionai /laion2b-23ish-1216px Overview This is a subset of https://huggingface.co/datasets/laion/laion2B-en-aesthetic, selected for aspect ratio, and with better captioning. Approximate image count is around 250k. 23ish, 1216px I picked out the images that are portrait aspect ratio of 2:3, or a little wider (Because images that are a little too wide, can be safely cropped narrower) I also picked a minimum height of 1216 pixels, because that is what 1024x1024 pixelcount converted to 2:3 looks like.… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-23ish-1216px.image100K<n<1M0 likes5 downloads2y agoHugging Face21opendiffusionai /cc12m-xlsdA quick upload of my current experimental dataset for XLSD. Approximately 350k images, ranging between 2mp and 4mp in size. Downloaded image disk space is somewhere under 400G An attempt has been made via AI to remove black and white images, non-real images, and watermarks, etc. So this is basically a "realistic photo" image set, with aspect ratios ranging from 3:2 to 2:3. We do not include 16:9 here, since I figured that would be a waste of information for a general purpose SD unet. This… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-xlsd.image100K<n<1M1 likes3 downloads2y agoHugging Face22opendiffusionai /laion2b-43ish-1000px Overview This is a subset of https://huggingface.co/datasets/laion/laion2B-en-aesthetic, selected for aspect ratio. It is a GENERAL CASE dataset. I attempted to filter out non-realistic images. This is intended to be a "real world" dataset. Approximate image count in this dataset is around 170k. On-disk size is around 185G. This is NOT individually human filtered. Almost no culling. That being said, the images are still relatively high quality! Approximately 4:3ish aspect ratio… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-43ish-1000px.image100K<n<1M1 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.