datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sdxl_10_reg
Stable Diffusion XL 1.0 Regularization Images
Note: All of these images were generated without the refiner. These are sdxl 1.0 base only.
This is some of my SDXL 1.0 regularization images generated with various prompts that are useful for regularization images or other specialized training. (color augmentation, bluring, shapening, etc). I will attempt to add more as I go along with various categories.
Each image has a corrisponding txt file with the prompt used to generate it as… See the full description on the dataset page: https://huggingface.co/datasets/ostris/sdxl_10_reg.sdxl_images_easy_prompts-artists-seed1sdxl-qwen-phase0
SDXL–Qwen Phase-0 dataset
Purpose-built training set for AbstractPhil/geolip-sdxl-aleph.
Each row pairs a Qwen-Image-Lightning render with the caption that produced it and an
encoder-invariant geometric "aleph" address derived from the caption's bytes. It exists to
retrain SDXL (which stays the base model) around a new text encoder (Qwen in place of
CLIP-G) under a rectified-flow objective: the render is the flow-matching target, and the
student learns to reproduce it from the… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/sdxl-qwen-phase0.danbooru2024-latents-sdxl-1ktar
Danbooru 2024 SDXL VAE latents in 1k tar
Dedicated dataset to align deepghs/danbooru2024-webp-4Mpixel. "4MP-Focus" for average raw image resolution.
Latents are ARB with maximum size of 1024x1024 as the recommended setting in kohyas. Major reason is to make sure I can finetune with RTX 3090. VRAM usage will raise drastically after 1024.
Generated from prepare_buckets_latents_v2.py, modified from prepare_buckets_latents.py.
Used for kohya-ss/sd-scripts. In theory it may replace… See the full description on the dataset page: https://huggingface.co/datasets/6DammK9/danbooru2024-latents-sdxl-1ktar.sdxl-1024-100kgenai-bench-sdxl-latents
GenAI-Bench SDXL with Latent Trajectories
SDXL generations for the GenAI-Bench prompts, paired with
the per-step decoded-latent trajectory of each generation. This dataset was created to evaluate NoisyCLIP.
For every prompt, 10 images were generated (1600 prompts → 16,000 generations). Each generation provides:
the full-resolution final image,
the 50-step denoising trajectory (each step's latent decoded to a small preview image), and
a CLIP-FlanT5-XXL VQAScore measuring… See the full description on the dataset page: https://huggingface.co/datasets/asiimo/genai-bench-sdxl-latents.sdxl-pickapicDanbooru-Top1000-Latents-SDXLsdxl-turbo-rtmesdxl-coco-filtered-fp32sdxl_images_sb_prompts-multi_artist-seed1image_generation_prompts_SDXL
Dataset Card for "image_generation_prompts_SDXL"
More Information needed
noun-attribute-sdxl-imagessdxl_laion_3parti-prompts-sdxl-1.0
Dataset Card for "parti-promtps-sdxl-1.0"
More Information needed
SDXL_Data_Curriculum_Step3SDXL_Data_Curriculum_Step2sdxl-datasetsdxl_laion_5_1style-content-grid-SDXL
Style Content Grid SDXL
Dataset Structure
The dataset contains 1738 images of resolution 1024x1024, generated by Stable Diffusion XL (sd_xl_base_1.0 with model hash 31e35c80fc). They were all generated in lllyasviel/stable-diffusion-webui-forge,
with the following positive and negative prompts:
Positive prompt: <style> of a <content>, \n masterpiece, best quality, high quality,
Negative prompt: (worst quality, low quality, normal quality),
with the following… See the full description on the dataset page: https://huggingface.co/datasets/yuxi-liu-wired/style-content-grid-SDXL.sdxl_laion_8_1sdxl-pickapic-fp32sdxl_laion_6_1sdxl_laion_7_0SDXL_Data_CSVformat_txtsdxl-qwen-phase1-cache
SDXL + Qwen3.5 Phase-1 Conditioning Cache
Precomputed, frozen-encoder conditioning for an SDXL + Qwen rectified-flow finetune: SDXL VAE latents, SDXL CLIP-L/CLIP-G text features, Qwen3.5 pooled and full-sequence text features, and geolip aleph addresses — one row per source image/caption pair, sharded so the build survives ephemeral compute and the result is reusable across runs and projects.
It exists so the expensive encode pass is paid once. Every array a downstream trainer… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/sdxl-qwen-phase1-cache.geolip-sdxl-fid-scoringautotrain-sdxl-vintage-face-style-lorasdxl-generated-10k
SDXL Generated Images Dataset (10,000 images)
This dataset contains 10,000 AI-generated images created with Stable Diffusion XL for training an AI image detector.
Dataset Details
Model: Stable Diffusion XL Base 1.0
Total Images: 10,000
Resolution: 1024×1024 pixels
Format: JPEG (quality 95)
Inference Steps: 10
Guidance Scale: 7.0
Random Seeds: Unique per image for maximum diversity
Generation Date: 2025-12-30
Prompt Diversity
Images generated with diverse… See the full description on the dataset page: https://huggingface.co/datasets/ash12321/sdxl-generated-10k.2M_fashionable_girl_SDXL_refiner_prompts
Dataset Card for "2M_fashionable_girl_SDXL_refiner_prompts"
More Information needed
