yuanchenyang/imagenet-256-flux2-vae-latents
ImageNet-256 FLUX.2 VAE Latents Pre-computed deterministic, model-facing encodings from the FLUX.2 VAE (black-forest-labs/FLUX.2-dev) for the full ImageNet-1K training set at 256x256 resolution, stored as Parquet shards. Each example includes latents for both the original and horizontally flipped image, enabling flip augmentation without re-encoding at training time. Dataset Description Each example contains: Column Shape Stored type Description… See the full description on the dataset page: https://huggingface.co/datasets/yuanchenyang/imagenet-256-flux2-vae-latents.
ImageNet-256 FLUX.2 VAE Latents
Pre-computed deterministic, model-facing encodings from the FLUX.2 VAE (`black-forest-labs/FLUX.2-dev`) for the full ImageNet-1K training set at 256x256 resolution, stored as Parquet shards. Each example includes latents for both the original and horizontally flipped image, enabling flip augmentation without re-encoding at training time.
Dataset Description
Each example contains:
- Number of examples: 1,281,167 (full ImageNet-1K train split)
- Raw VAE latent:
(32, 32, 32)(8x spatial downsampling) - Model-facing latent:
(128, 16, 16)after 2x2 patchification - Precision: float16 storage; loaders may promote values to float32
- Parquet shards: 641
- Total Parquet size: 194,204,502,807 bytes
The column names match `yuanchenyang/imagenet-256-sd-vae-ft-mse-latents` for loader compatibility. Despite the latent_mean name, values are not raw, unnormalized posterior means. They are ready to use as FLUX.2 model inputs.
Creation
Images were center-cropped and resized to 256x256 using the Dhariwal (ADM) cropping method, then normalized to [-1, 1] before encoding.
For each original and horizontally flipped image:
- The FLUX.2 VAE posterior mode was computed, producing
(32, 32, 32). - Each 2x2 spatial neighborhood was packed into channels, producing
(128, 16, 16). - The packed latent was normalized using the VAE checkpoint's batch-normalization running mean, running variance, and
eps=1e-4. - The model-facing encoding was cast to float16 for storage.
The VAE checkpoint SHA-256 used for this export is d64f3a68e1cc4f9f4e29b6e0da38a0204fe9a49f2d4053f0ec1fa1ca02f9c4b5.
Reconstruction Validation
A class-balanced 1,000-image pilot was validated for both original and flipped images (2,000 reconstructions total). Float16 storage was compared with the same model-facing encodings retained in float32:
All 641 exported shards were subsequently checked for row count, float16 schema, and label-order integrity.
Usage
import torch
from datasets import load_dataset
ds = load_dataset(
"YOUR_NAMESPACE/imagenet-256-flux2-vae-latents",
split="train",
)
example = ds[0]
latent = torch.tensor(example["latent_mean"], dtype=torch.float32)
latent_flip = torch.tensor(example["latent_mean_flip"], dtype=torch.float32)
label = example["label"]The latent values are already normalized and model-facing. Do not apply an additional VAE scaling factor or shift.
Intended Use
Training latent diffusion and flow-matching models on ImageNet-256 without running the FLUX.2 VAE encoder during training.
Source Code
The export and validation scripts are available in `chenyang-tri/diffusion_baseline`:
flux2_vae.pyprepare_flux2_dataset.pyexport_flux2_dataset.py
License
The source images and derived latent representations inherit the ImageNet terms of access. The FLUX.2 VAE weights are distributed under the FLUX Non-Commercial License. Users are responsible for complying with both sets of terms.
