datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imagenet-256-flux2-vae-latents
ImageNet-256 FLUX.2 VAE Latents
Pre-computed deterministic, model-facing encodings from the
FLUX.2 VAE (black-forest-labs/FLUX.2-dev)
for the full ImageNet-1K training set at 256x256 resolution, stored as Parquet
shards. Each example includes latents for both the original and horizontally
flipped image, enabling flip augmentation without re-encoding at training time.
Dataset Description
Each example contains:
Column
Shape
Stored type
Description… See the full description on the dataset page: https://huggingface.co/datasets/yuanchenyang/imagenet-256-flux2-vae-latents.danbooru2024-latents-sdxl-1ktar
Danbooru 2024 SDXL VAE latents in 1k tar
Dedicated dataset to align deepghs/danbooru2024-webp-4Mpixel. "4MP-Focus" for average raw image resolution.
Latents are ARB with maximum size of 1024x1024 as the recommended setting in kohyas. Major reason is to make sure I can finetune with RTX 3090. VRAM usage will raise drastically after 1024.
Generated from prepare_buckets_latents_v2.py, modified from prepare_buckets_latents.py.
Used for kohya-ss/sd-scripts. In theory it may replace… See the full description on the dataset page: https://huggingface.co/datasets/6DammK9/danbooru2024-latents-sdxl-1ktar.imagenet-256-sd-vae-ft-mse-latents
ImageNet-256 SD-VAE-ft-MSE Latents
Pre-computed posterior means (no variance/std) from the Stable Diffusion VAE (stabilityai/sd-vae-ft-mse) for the full ImageNet-1K training set at 256×256 resolution, stored as Parquet shards. Each example includes latents for both the original and horizontally flipped image, enabling flip augmentation without re-encoding at training time.
Dataset Description
Each example contains:
Column
Shape
Type
Description
latent_mean… See the full description on the dataset page: https://huggingface.co/datasets/yuanchenyang/imagenet-256-sd-vae-ft-mse-latents.e621_2024-latents-sdxl-1ktar
E621 2024 SDXL VAE latents in 1k tar
Dedicated dataset to align both NebulaeWis/e621-2024-webp-4Mpixel and deepghs/e621_newest-webp-4Mpixel. "4MP-Focus" for average raw image resolution.
Latents are ARB with maximum size of 1024x1024 as the recommended setting in kohyas. Major reason is to make sure I can finetune with RTX 3090. VRAM usage will raise drastically after 1024.
Generated from prepare_buckets_latents_v2.py, modified from prepare_buckets_latents.py.
Used for… See the full description on the dataset page: https://huggingface.co/datasets/6DammK9/e621_2024-latents-sdxl-1ktar.
