CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Gustavosta /Stable-Diffusion-Prompts Stable Diffusion Dataset This is a set of about 80,000 prompts filtered and extracted from the image finder for Stable Diffusion: "Lexica.art". It was a little difficult to extract the data, since the search engine still doesn't have a public API without being protected by cloudflare. If you want to test the model with a demo, you can go to: "spaces/Gustavosta/MagicPrompt-Stable-Diffusion". If you want to see the model, go to: "Gustavosta/MagicPrompt-Stable-Diffusion". text10K<n<100K527 likes5.5k downloads4y agoHugging Face02hanamizuki-ai /stable-diffusion-v1-5-glazed Dataset Card for Stable Diffusion v1.5 Glazed Samples Dataset Description Dataset Summary This dataset contains image samples originally generated by runwayml/stable-diffusion-v1-5 and subsequently processed by Glaze tool. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/stable-diffusion-v1-5-glazed.imageimage-classification100K<n<1M3 likes3.1k downloads3y agoHugging Face03AbstractPhil /diffusion-pretrain-set-ft1 diffusion-pretrain-set-ft1 A multi-source image-caption pretraining dataset assembled from ten upstream sources via a uniform ingest pipeline. Designed for a full pretrain or finetune pipeline meant to curate for any major diffusion model preliminary, with the sole intent to create a more powerful baseline preliminary train and a baseline for synthesizing images to train the next generation of the VLM model. This is a lot like the snake eating it's own tail, so it must be… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/diffusion-pretrain-set-ft1.image1M<n<10M2 likes1.9k downloads3mo agoHugging Face04DenisKochetov /Diffusion4D-Animated-Raw Diffusion4D Animated Assets This dataset provides animated 3D assets referenced by Diffusion4D and Objaverse-XL in a directly browsable format. The default split contains 67,988 rows. Each row includes metadata, a preview image, and a short preview video so that assets can be inspected in the Hugging Face Data Studio without first downloading the original 3D file. The repository also mirrors available raw assets and keeps their original source links and hashes. The current… See the full description on the dataset page: https://huggingface.co/datasets/DenisKochetov/Diffusion4D-Animated-Raw.3d10K<n<100K4 likes1.4k downloads3mo agoHugging Face05gmongaras /Stable_Diffusion_3_RecaptionThis dataset is the one specified in the stable diffusion 3 paper which is composed of the ImageNet dataset and the CC12M dataset. I used the ImageNet 2012 train/val data and captioned it as specified in the paper: "a photo of a 〈class name〉" (note all ids are 999,999,999) CC12M is a dataset with 12 million images created in 2021. Unfortunately the downloader provided by Google has many broken links and the download takes forever. However, some people in the community publicized the dataset.… See the full description on the dataset page: https://huggingface.co/datasets/gmongaras/Stable_Diffusion_3_Recaption.image10M<n<100M5 likes1.4k downloads2y agoHugging Face06rbeauchamp /diffusion_db_dedupe_from50k_train Dataset Card for "diffusion_db_dedupe_from50k_train" More Information needed image10K<n<100K1 likes1k downloads3y agoHugging Face07AbstractPhil /diffusion-pretrain-set-ft1-1024 diffusion-pretrain-set-ft1-1024 1024px (2x) upscale of AbstractPhil/diffusion-pretrain-set-ft1. WARNING MUCH OF THIS DATA WAS MODEL UPSCALED USING RAPID UPSCALERS. THIS IS NOT CONSISTENTLY HIGH FIDELITY NOR IS IT EVEN CLOSE TO FAIR FIDELITY AT TIMES. PLEASE use this ONLY for pretraining, new concepts, and simple design purposes ONLY. HEAVILY PRUNE FOR FINETUNING. Thank you, good luck my friends. Details Model: realesr-general-x4v3 (SRVGG Compact… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/diffusion-pretrain-set-ft1-1024.image1M<n<10M0 likes966 downloads4mo agoHugging Face08fffffchopin /DiffusionDream_DatasetThis is the dataset of the diffusion dream dataset. The dataset contains the following columns: info: A string describing the action taken in the frame keyword: A string describing the keyword of the action action: A string describing the action taken in the frame current_frame: The current frame of the video previous_frame_1: The frame before the current frame previous_frame_2: The frame before the previous frame previous_frame_3: The frame before the previous frame previous_frame_4: The… See the full description on the dataset page: https://huggingface.co/datasets/fffffchopin/DiffusionDream_Dataset.image100K<n<1M1 likes611 downloads1y agoHugging Face09giakoupg /diffusionprint_dataset DiffusionPrint Patch Dataset A dataset of 64x64 image patches for contrastive learning of diffusion-based inpainting forensics. Each patch comes from either a real image or an AI-inpainted region generated by one of three diffusion models. Columns Column Type Description image bytes (PNG) Lossless 64x64 RGB patch master_index int64 Row index in the original memmap archive patch_path string Original relative path of the patch category string real… See the full description on the dataset page: https://huggingface.co/datasets/giakoupg/diffusionprint_dataset.textimage-classification1M<n<10M0 likes603 downloads3mo agoHugging Face10jtatman /stable-diffusion-prompts-stats-full-uncensoredimage100K<n<1M152 likes547 downloads2y agoHugging Face11gzzyyxy /layout_diffusion_hypersimThis repository contains the data for SceneCraft: Layout-Guided 3D Scene Generation. Project page: https://orangesodahub.github.io/SceneCraft Code: https://github.com/OrangeSodahub/SceneCraft imagetext-to-3d10K<n<100K1 likes546 downloads1y agoHugging Face12svjack /diffusiondb_2m_random_50k Dataset Card for "diffusiondb_2m_random_50k" More Information needed image10K<n<100K0 likes516 downloads4y agoHugging Face13LGirrbach /person-centric-images-stable-diffusion-v1-1image100K<n<1M0 likes494 downloads1y agoHugging Face14yuwan0 /lexica-stable-diffusion-v1-5 Stable Diffusion Dataset This is a set of about 80,000 Image-Prompt pairs generated by stable-diffusion-v1-5. The Prompts come from dataset Stable-Diffusion-Prompts which filtered and extracted from the image finder for Stable Diffusion: "Lexica.art". image10K<n<100K4 likes369 downloads3y agoHugging Face15rdoshi21 /1m2r-real-diffusion2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "kinova-arx", "total_episodes": 72, "total_frames": 32350, "total_tasks": 1, "total_videos": 0, "total_chunks": 1, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:72" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-real-diffusion2.imagerobotics10K<n<100K0 likes346 downloads8mo agoHugging Face16teticio /audio-diffusion-1024Over 20,000 256x256 mel spectrograms of 5 second samples of music from my Spotify liked playlist. The code to convert from audio to spectrogram and vice versa can be found in https://github.com/teticio/audio-diffusion along with scripts to train and run inference using De-noising Diffusion Probabilistic Models. x_res = 1024 y_res = 1024 sample_rate = 44100 n_fft = 2048 hop_length = 512 imageimage-to-image10K<n<100K0 likes328 downloads4y agoHugging Face17cyttic /diffusionpen-hebrew-handwriting DiffusionPen Hebrew Handwriting A large synthetic dataset of Hebrew handwritten text lines with ground-truth transcriptions, for training and evaluating handwritten text recognition (HTR / OCR) models. Every image is a single line of right-to-left Hebrew handwriting synthesized by DiffusionPen — a style-conditioned latent-diffusion handwriting generator — in one of 491 distinct writer styles, and quality-filtered by an independent OCR pass. 149,952 line images, 491 writer… See the full description on the dataset page: https://huggingface.co/datasets/cyttic/diffusionpen-hebrew-handwriting.imageimage-to-text100K<n<1M1 likes320 downloads3mo agoHugging Face18thefcraft /civitai-stable-diffusion-337k How to Use from datasets import load_dataset dataset = load_dataset("thefcraft/civitai-stable-diffusion-337k") print(dataset['train'][0]) download images download zip files from images dir https://huggingface.co/datasets/thefcraft/civitai-stable-diffusion-337k/tree/main/images it contains some images with id from zipfile import ZipFile with ZipFile("filename.zip", 'r') as zObject: zObject.extractall() Dataset Summary GitHub URL:-… See the full description on the dataset page: https://huggingface.co/datasets/thefcraft/civitai-stable-diffusion-337k.image100K<n<1M43 likes317 downloads2y agoHugging Face19punwaiw /DiffusionJockey Dataset Card for "DiffusionJockey" More Information needed image1K<n<10K0 likes300 downloads3y agoHugging Face20ririye /Benchmark-Images-for-Stable-Diffusion-Biastext1K<n<10K0 likes300 downloads2y agoHugging Face21bitmind /bm-subnet-stable-diffusion-xl-base-1.0image10K<n<100K0 likes291 downloads2y agoHugging Face22myradeng /diffusion_db_dedup_from50k_train_v2 Dataset Card for "diffusion_db_dedup_from50k_train_v2" More Information needed image1K<n<10K0 likes261 downloads3y agoHugging Face23whosouravsharma /text-to-image-diffusiondb-2M DiffusionDB text-to-image subset A cleaned, safety-filtered image-prompt dataset for training a text-to-image model, built from DiffusionDB. Built on Hugging Face Jobs directly from poloclub/diffusiondb. It covers part_id 1-20 (20,000 source images) before filtering. The same content is also kept on the 20k-subset branch. Load it with: load_dataset("whosouravsharma/text-to-image-diffusiondb-2M") Note on the repo name: despite "2M" in the name, this is a small slice of… See the full description on the dataset page: https://huggingface.co/datasets/whosouravsharma/text-to-image-diffusiondb-2M.imagetext-to-image10K<n<100K0 likes202 downloads1mo agoHugging Face24svjack /diffusiondb_random_10k_zh_v1 Dataset Card for "diffusiondb_random_10k_zh_v1" svjack/diffusiondb_random_10k_zh_v1 is a dataset that random sample 10k English samples from diffusiondb and use NMT translate them into Chinese with some corrections. it used to train stable diffusion models in svjack/Stable-Diffusion-FineTuned-zh-v0 svjack/Stable-Diffusion-FineTuned-zh-v1 svjack/Stable-Diffusion-FineTuned-zh-v2 And is the data support of https://github.com/svjack/Stable-Diffusion-Chinese-Extend which is a fine tune… See the full description on the dataset page: https://huggingface.co/datasets/svjack/diffusiondb_random_10k_zh_v1.image1K<n<10K3 likes167 downloads4y agoHugging Face25svjack /diffusiondb_random_10k Dataset Card for "diffusiondb_random_10k" More Information needed image10K<n<100K0 likes164 downloads4y agoHugging Face260xJustin /Dungeons-and-DiffusionThis is the dataset! Not the .ckpt trained model - the model is located here: https://huggingface.co/0xJustin/Dungeons-and-Diffusion/tree/main The newest version has manually captioned races and classes, and the model is trained with EveryDream. 30 images each of: aarakocra, aasimar, air_genasi, centaur, dragonborn, drow, dwarf, earth_genasi, elf, firbolg, fire_genasi, gith, gnome, goblin, goliath, halfling, human, illithid, kenku, kobold, lizardfolk, minotaur, orc, tabaxi, thrikreen, tiefling… See the full description on the dataset page: https://huggingface.co/datasets/0xJustin/Dungeons-and-Diffusion.image1K<n<10K72 likes154 downloads3y agoHugging Face27teticio /audio-diffusion-512Over 20,000 512x512 mel spectrograms of 5 second samples of music from my Spotify liked playlist. The code to convert from audio to spectrogram and vice versa can be found in https://github.com/teticio/audio-diffusion along with scripts to train and run inference using De-noising Diffusion Probabilistic Models. x_res = 512 y_res = 512 sample_rate = 22050 n_fft = 2048 hop_length = 512 imageimage-to-image10K<n<100K2 likes152 downloads3y agoHugging Face28jikaixuan /diffusion_sfttext100K<n<1M0 likes151 downloads3y agoHugging Face29cyttic /diffusionpen-hebrew-handwriting-cer0 DiffusionPen Hebrew Handwriting — CER=0 clean subset The highest-fidelity slice of DiffusionPen Hebrew Handwriting: only the 33,082 line images that an independent Hebrew TrOCR read back with an exact match (character error rate = 0.0). Built to test whether a smaller, label-clean set trains a better recognizer than the full (noisier) 150k set. 33,082 line images, 491 writer styles Splits (writer-independent, style-disjoint): train 26,607 / validation 3,187 / test 3,288 Every… See the full description on the dataset page: https://huggingface.co/datasets/cyttic/diffusionpen-hebrew-handwriting-cer0.imageimage-to-text10K<n<100K0 likes144 downloads3mo agoHugging Face30bayes-group-diffusion /wikipedia-bloom Dataset Card for "wikipedia-bloom" More Information needed text1M<n<10M0 likes142 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.