open-diffusion
pexels-photos-janpf
Images:
There are approximately 130K images, borrowed from pexels.com.
Thanks to those folks for curating a wonderful resource.
There are millions more images on pexels. These particular ones were selected by
the list of urls at https://github.com/janpf/self-supervised-multi-task-aesthetic-pretraining/blob/main/dataset/urls.txt .
The filenames are based on the md5 hash of each image.
Download From here or from pexels.com: You choose
For those people who like… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/pexels-photos-janpf.cc12m-cleaned
CC12m-cleaned
This dataset builds on two others: The Conceptual Captions 12million dataset, which lead to the LLaVa captioned subset done by
CaptionEmporium
(The latter is the same set, but swaps out the (Conceptual Captions 12million) often-useless alt-text captioning for decent ones_
I have then used the llava captions as a base, and used the detailed descrptions to filter out
images with things like watermarks, artist signatures, etc.
I have also manually thrown out all… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned.eval_diffusion-v2-augThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/open-cloth/eval_diffusion-v2-aug.eval_diffusion-cloth-LR-asyncThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 15,
"total_frames": 23605,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:15"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/open-cloth/eval_diffusion-cloth-LR-async.pexels-janpf-sharp
Overview
This is a strict subset of
https://huggingface.co/datasets/opendiffusionai/pexels-photos-janpf
I have attempted to throw out all shots with heavy bokeh. (so, this is the "sharp" focus dataset)
I have also attempted to throw out all the black-and-white photos.
I decided to create a whole "new" dataset, rather than creating a set filter as I
have done previously, because I think this dataset may become my new "base" dataset.
So I will most likely focus my refining… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/pexels-janpf-sharp.cc12m-a_woman
Description
This dataset is a convenience subset of https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned/
I did a quick grep for "A woman", and then HAND-CURATED the results.
That means I threw out anything with watermarks, site branding, or pretty much anything else I deemed would
get in the way of ML training.
I also only chose images that had clear, sharp camera focus on the main subject. So these are high-quality images.
At present, I have only done a few thousand.
I… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-a_woman.
