CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01opendiffusionai /pexels-photos-janpf Images: There are approximately 130K images, borrowed from pexels.com. Thanks to those folks for curating a wonderful resource. There are millions more images on pexels. These particular ones were selected by the list of urls at https://github.com/janpf/self-supervised-multi-task-aesthetic-pretraining/blob/main/dataset/urls.txt . The filenames are based on the md5 hash of each image. Download From here or from pexels.com: You choose For those people who like… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/pexels-photos-janpf.text-to-image100K<n<1M45 likes666 downloads8mo agoHugging Face02opendiffusionai /cc12m-cleaned CC12m-cleaned This dataset builds on two others: The Conceptual Captions 12million dataset, which lead to the LLaVa captioned subset done by CaptionEmporium (The latter is the same set, but swaps out the (Conceptual Captions 12million) often-useless alt-text captioning for decent ones_ I have then used the llava captions as a base, and used the detailed descrptions to filter out images with things like watermarks, artist signatures, etc. I have also manually thrown out all… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned.imagetext-to-image1M<n<10M13 likes269 downloads2y agoHugging Face03open-cloth /eval_diffusion-v2-augThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "observation.state": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/open-cloth/eval_diffusion-v2-aug.tabularrobotics10K<n<100K0 likes232 downloads4mo agoHugging Face04open-cloth /eval_diffusion-cloth-LR-asyncThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 15, "total_frames": 23605, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:15" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/open-cloth/eval_diffusion-cloth-LR-async.tabularrobotics10K<n<100K0 likes180 downloads4mo agoHugging Face05opendiffusionai /pexels-janpf-sharp Overview This is a strict subset of https://huggingface.co/datasets/opendiffusionai/pexels-photos-janpf I have attempted to throw out all shots with heavy bokeh. (so, this is the "sharp" focus dataset) I have also attempted to throw out all the black-and-white photos. I decided to create a whole "new" dataset, rather than creating a set filter as I have done previously, because I think this dataset may become my new "base" dataset. So I will most likely focus my refining… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/pexels-janpf-sharp.text10K<n<100K3 likes106 downloads2mo agoHugging Face06opendiffusionai /cc12m-a_woman Description This dataset is a convenience subset of https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned/ I did a quick grep for "A woman", and then HAND-CURATED the results. That means I threw out anything with watermarks, site branding, or pretty much anything else I deemed would get in the way of ML training. I also only chose images that had clear, sharp camera focus on the main subject. So these are high-quality images. At present, I have only done a few thousand. I… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-a_woman.image10K<n<100K27 likes79 downloads2y agoHugging Face07opendiffusionai /laion2b-23ish-woman-solo Overview All images have a woman in them, solo, at APPROXIMATELY 2:3 aspect ratio. These images are HUMAN CURATED. I have personally gone through every one at least once. Additionally, there are no visible watermarks, the quality and focus are good, and it should not be confusing for AI training There should be a little over 15k images here. Note that there is a wide variety of body sizes, from size 0, to perhaps size 18 There are also THREE choices of captions: the really bad "alt… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-23ish-woman-solo.image10K<n<100K19 likes70 downloads1y agoHugging Face08opendiffusionai /laion2b-en-aesthetic-square-human Overview This dataset is a HAND-CURATED version of our laion2b-en-aesthetic-square-cleaned dataset. It has at least "a man" or "a woman" in it. Additionally, there are no visible watermarks, the quality and focus are good, and it should not be confusing for AI training There should be a little over 8k images here. Details It consists of an initial extraction of all images that had "a man" or "a woman" in the moondream caption. I then filtered out all "statue" or… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-en-aesthetic-square-human.image1K<n<10K8 likes66 downloads2y agoHugging Face09opendiffusionai /cc12m-4mp-realistic Overview This is a hand-selected subset of our larger attempts to filter the well known CC12M dataset. This one focuses on large (4 megapixels) images that are real world, high quality images, and the captioning specifically matches either "A man" or "A woman". Note that I did not have the diskspace/time to go through the ENTIRE set. It was perhaps only from the first 2 million of our CC12M-cleaned subset. If an effort were made to go through the entire 4mp image set, there might be… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-4mp-realistic.imagetext-to-image10K<n<100K22 likes53 downloads2y agoHugging Face10opendiffusionai /cc12m-2mp-realistic Overview A subset of the "CC12m" dataset. Size varying between 2mp <= x < 4mp. Anything actually 4mp or larger will be in one of our "4mp" datasets. This dataset is created for if you basically need more images and dont mind a little less quality. Quality I have filtered out as many watermarks, etc. as possible using AI models. I have also thrown out stupid black-and-white photos, because they poison normal image prompting. This is NOT HAND CURATED, unlike some… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-2mp-realistic.image100K<n<1M4 likes42 downloads2y agoHugging Face11opendiffusionai /laion2b-en-aesthetic-square-cleaned Overview A subset of our opendiffusionai/laion2b-en-aesthetic-square, which is itself a subset of the widely known "Laion2b-en-aesthetic" dataset. However, the original had only the website alt-tags for captions. I have added decent AI captioning, via the "moondream2" model. Additionally, I have stripped out a bunch of watermarked junk, and weeded out 20k duplicate images. When I use it for training, I will be trimming out additional things like non-realistic images, painting… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-en-aesthetic-square-cleaned.image100K<n<1M11 likes38 downloads8mo agoHugging Face12open-cloth /eval_diffusion-clothThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so_follower", "total_episodes": 2, "total_frames": 2886, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/open-cloth/eval_diffusion-cloth.tabularrobotics1K<n<10K0 likes35 downloads4mo agoHugging Face13opendiffusionai /cc12m-1mp_plus-realistic cc12m-1mp_plus-realistic A filtering down of the full CC12M dataset, to have the following characteristics: At least 1024x1024 pixels in size "Realistic". No paintings, digital art, monochrome, or surreal stuff. Also discard multi-image as much as possible Ideally, no signed or watermarked images. (but there will certainly be some left) Captions The caption types available are a bit different from some of our other ones. Currently available are: caption_llava… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-1mp_plus-realistic.image100K<n<1M4 likes31 downloads10mo agoHugging Face14opendiffusionai /pexels-woman-solo Overview Around 8,000 hand-selected high-resolution 4k images of "a woman", suitable for both training, and "pre-training" of AI models. Images are all real-world realism based. Background I have been having difficulty training my early-stage txt2img model with GOOD, HIGH-RES images of what "a woman" is. Up until now, I have just been throwing a large number of random high-res images with "woman" in the auto-captioned details. NOW, however, I have hand-selected a bunch of… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/pexels-woman-solo.text-to-image1K<n<10K3 likes27 downloads2y agoHugging Face15opendiffusionai /woman-wipThis dataset is just a collaborative area for now, to trim down a candidate set of images image1 likes26 downloads2y agoHugging Face16opendiffusionai /pexels-woman-croppable pexels-woman-croppable An extract from our larger "pexels 130k images" set. Around 6500 images. Useful for training text-to-image models. But we already have a bunch of subsets, why another one? This is 6k images that, while not originally square cropped, are HAND-SELECTED to be square-crop clean. The provided crawl.sh util script will handle automatically downloading and cropping them from pexels.com Why Square-Crop When training a model, you must have all images… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/pexels-woman-croppable.imagetext-to-image1K<n<10K0 likes24 downloads8mo agoHugging Face17opendiffusionai /laion2b-aesthetic-squareish-captionsThis dataset contains image captions generated from LAION2B-en-aesthetic-square. We started with ~300K images after size filtering (2.5k max w/h), a portion of the images were skipped due to inaccessible URLs. The captions were generated over ~30 hours using Qwen3-VL-30B-A3B-Instruct on 1xH100 running SGLang with the prompt Describe the content of the provided image in detail, in plaintext. Do not make assumptions. Do not use special formatting. Avoid purple prose. Total samples: 209,141… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-aesthetic-squareish-captions.image100K<n<1M1 likes22 downloads10mo agoHugging Face18opendiffusionai /cc12m-4mp Parent dataset https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned Contents I had uploaded a prior version of this "4mp" set, that was just partial. Now I have grabbed all of them available as of 2024 December. This latest upload references ALL files in the "CC12M" dataset, that are at least 4 megapixel in size, and are not one of the known "stock images" copyrighted sites. (and were available this month) I have done some additional filtering out of any… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-4mp.image100K<n<1M3 likes21 downloads2y agoHugging Face19opendiffusionai /laion2b-mixed-1024px-human Overview This dataset is a selective merge of some other of our datasets. Mainly, I pulled human-centric, real-world photos from the following datasets: opendiffusionai/laion2b-45ish-1120px opendiffusionai/laion2b-squareish-1024px opendiffusionai/laion2b-23ish-1216px As such, it is "mixed" aspect ratio. The very smallest height ones are from our "squarish" set, so are at least 1024px tall. However, the other ones with longer rations, have an appropriately longer minimum… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-mixed-1024px-human.image10K<n<100K3 likes21 downloads2y agoHugging Face20opendiffusionai /laion2b-squareish-1536px Overview This dataset is a subset of laion2b-en-aesthetic dataset. Its purpose is for general-case (photographic) model training. Its aspect ratios are (1 <= ratio < 5/4 ) so that images can be auto-cropped to square safely. Additionally, all images are at least 1536 pixels tall. This is because I wanted a dataset that would be very high quality to use for models that are 768x768 There should be close to 80k 8k images here. Note that you can CHOOSE between "moondream"… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-squareish-1536px.image10K<n<100K3 likes19 downloads2y agoHugging Face21opendiffusionai /laion2b-squareish-1024px Overview This dataset is a subset of laion2b-en-aesthetic dataset. Its purpose is for general-case (photographic) model training. Its aspect ratios are (1 <= ratio < 5/4 ) so that images can be auto-cropped to square safely. Additionally, all images are at least 1024 pixels tall. There should be close to 250k images here. Total on-disk size is approximately 118G We have a very similar dataset that is 1536px in size. However, that is only 80k in size, so if you need waaay more… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-squareish-1024px.image100K<n<1M1 likes19 downloads2y agoHugging Face22opendiffusionai /laion2b-45ish-1120px Overview This is a subset of https://huggingface.co/datasets/laion/laion2B-en-aesthetic, selected for aspect ratio, and with better captioning. It is a GENERAL CASE dataset. Percentages of images with humans in it, is approximately only 30% I also filtered out all non-realistic images. This is intended to be a "real world" dataset. Approximate image count in this dataset is around 80k. On-disk size is around 45G. This is NOT individually human filtered. Batch culling only.… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-45ish-1120px.image10K<n<100K3 likes17 downloads2y agoHugging Face23diffusion-guidance-ku-hf /sdxl_open-image-preferences_0text1K<n<10K0 likes13 downloads2y agoHugging Face24opendiffusionai /laion2b-squareish General purpose dataset (realistic) This is a merging of our laion2b-squareish-1024px and laion2b-squareish-1536px datasets, WITH some additional filtering. It has a little de-duplication, and some extra watermark removal. It is MOSTLY realistic. I'm sharing this, because I am actively using it in my SD1.5 training experiments. I needed specifically square(ish) rather than our other nice, but mixed aspect-ratio datasets. Plus, I needed a LARGE one. So, here it is!… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-squareish.image100K<n<1M0 likes13 downloads1y agoHugging Face25opendiffusionai /cc12m-xlsd-512px What Assorted images from our CC12M filtered sets, centercropped and then resized to 512x512 pro-actively Why This probably wont be useful to people outside the org, but just in case... here you go! The wierd naming matches the MD5 checksum of the ORIGINAL FULL-SIZED image. Internally i sort all my images like this. So I can do filtering or autotagging on the mini-images, then directly apply it to the originals. image100K<n<1M0 likes9 downloads2y agoHugging Face26opendiffusionai /cc12m-small-squarish-simple Contents of dataset This is a "general purposes dataset", of images that are 512px to 1024px in size, if I recall correctly (in contrast to the "2mp" and "4mp" datasets) Also, they are "squarish", which means that, even if they are not precisely square, the image looks fine if you hard-crop it to force square. Additionally, images have been hand-culled to throw out anything I considered bad for AI training. Additionally, they were AI-culled to have a "simple photographic… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-small-squarish-simple.image10K<n<100K0 likes9 downloads1y agoHugging Face27opendiffusionai /laion2b-en-aesthetic-square Contents This is a pretty raw filter of https://huggingface.co/datasets/laion/laion2B-en-aesthetic I just filtered for "is image perfectly square, AND is image at least 1024x1024 pixels" Approximate image count is a bit over 300k Updated 2025/01/24 I just found out there are a bunch of watermarked sites in here. So much for aesthetically chosen :( So I filtered out a bunch of the "stock image" sites, just by looking at url strings. Update 2025/01/31… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-en-aesthetic-square.image100K<n<1M1 likes8 downloads2y agoHugging Face28opendiffusionai /cc12m-2mp-squareish Overview A subset of the "CC12m" dataset. Around 37k images, varying between 2mp <= x < 4mp. Anything actually 4mp or larger will be in one of our "4mp" datasets. This dataset is created for if you basically need more images and dont mind a little less quality. Why "Squareish"? I noticed that SD1.5 really, really likes square images to train on. These are mostly NOT EXACTLY SQUARE. However, most training programs should have an auto-crop capability. This dataset is of… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-2mp-squareish.image10K<n<100K0 likes6 downloads2y agoHugging Face29opendiffusionai /laion2b-34ish-1152pxDont use this yet. Work in progress. In theory you could, but its messy data. I uploaded it here so I could more easily transfer it off my home machine to a research machine, where I will then filter it for watermarks and give it better captions. Should take maybe a week. image100K<n<1M0 likes6 downloads10mo agoHugging Face30opendiffusionai /laion2b-23ish-1216px Overview This is a subset of https://huggingface.co/datasets/laion/laion2B-en-aesthetic, selected for aspect ratio, and with better captioning. Approximate image count is around 250k. 23ish, 1216px I picked out the images that are portrait aspect ratio of 2:3, or a little wider (Because images that are a little too wide, can be safely cropped narrower) I also picked a minimum height of 1216 pixels, because that is what 1024x1024 pixelcount converted to 2:3 looks like.… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-23ish-1216px.image100K<n<1M0 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.