datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
danbooru2024-latents-sdxl-1ktar
Danbooru 2024 SDXL VAE latents in 1k tar
Dedicated dataset to align deepghs/danbooru2024-webp-4Mpixel. "4MP-Focus" for average raw image resolution.
Latents are ARB with maximum size of 1024x1024 as the recommended setting in kohyas. Major reason is to make sure I can finetune with RTX 3090. VRAM usage will raise drastically after 1024.
Generated from prepare_buckets_latents_v2.py, modified from prepare_buckets_latents.py.
Used for kohya-ss/sd-scripts. In theory it may replace… See the full description on the dataset page: https://huggingface.co/datasets/6DammK9/danbooru2024-latents-sdxl-1ktar.sdxl-generated-10k
SDXL Generated Images Dataset (10,000 images)
This dataset contains 10,000 AI-generated images created with Stable Diffusion XL for training an AI image detector.
Dataset Details
Model: Stable Diffusion XL Base 1.0
Total Images: 10,000
Resolution: 1024×1024 pixels
Format: JPEG (quality 95)
Inference Steps: 10
Guidance Scale: 7.0
Random Seeds: Unique per image for maximum diversity
Generation Date: 2025-12-30
Prompt Diversity
Images generated with diverse… See the full description on the dataset page: https://huggingface.co/datasets/ash12321/sdxl-generated-10k.e621_2024-latents-sdxl-1ktar
E621 2024 SDXL VAE latents in 1k tar
Dedicated dataset to align both NebulaeWis/e621-2024-webp-4Mpixel and deepghs/e621_newest-webp-4Mpixel. "4MP-Focus" for average raw image resolution.
Latents are ARB with maximum size of 1024x1024 as the recommended setting in kohyas. Major reason is to make sure I can finetune with RTX 3090. VRAM usage will raise drastically after 1024.
Generated from prepare_buckets_latents_v2.py, modified from prepare_buckets_latents.py.
Used for… See the full description on the dataset page: https://huggingface.co/datasets/6DammK9/e621_2024-latents-sdxl-1ktar.sdxl-base-1-scm-corpus
Shamima/sdxl-base-1-scm-corpus
Synthetic image corpus generated with Stable Diffusion XL for studying the
Stereotype Content Model (SCM) structure of text-to-image latent space.
Images: 6,600
Categories: 66 occupation/identity groups
Prompt template: "A portrait of a [group], high quality."
Generator: SDXL base 1.0, DPM++ 2M Karras, 30 steps, CFG 7.0
Resolution: see image features
Fields
field
description
image
RGB JPEG
category
Group/occupation label… See the full description on the dataset page: https://huggingface.co/datasets/Shamima/sdxl-base-1-scm-corpus.social-media-robustness-sdxl-instantid
Social Media Robustness Benchmark: SDXL+InstantID Synthetic Face Detection
Version: v1.0.0 · License: CC BY-NC 4.0 (research evaluation only)
Detector accuracy on clean lab test sets does not predict in-the-wild performance. Social
platforms re-encode every uploaded image: platform-specific JPEG, resize, chroma subsampling,
metadata stripped. This benchmark lets detector authors and procurers measure robustness under
documented, paired, demographically-balanced conditions… See the full description on the dataset page: https://huggingface.co/datasets/danb21/social-media-robustness-sdxl-instantid.synthetic-face-sdxl-instantid-bench
Synthetic Face Detection Benchmark — SDXL+InstantID
Version: v1.0.0 · Build date: 2026-05-16 · Rows: 26492
Evaluation benchmark for synthetic-face detection under platform-realistic
conditions. Sampled to satisfy the ISO/IEC 19795 floor of 300 samples per
demographic subgroup across a 6×2 (skin tone × gender) cell grid. Not
training data; not licensed for commercial use.
See release.json for build provenance, manifest.csv for per-row
license attestation, LICENSES.csv for the… See the full description on the dataset page: https://huggingface.co/datasets/danb21/synthetic-face-sdxl-instantid-bench.
