dingshizhe/vgtae-minit2i-sft-latents
MiniT2I SFT-set latents encoded with VGT-AE (448px) The text-to-image SFT mixture, pre-encoded into the VGT-AE latent space. Companion to dingshizhe/vgtae-cc12m-latents, which covers the pretraining corpus; same encoder, same settings, same layout. These are not DC-AE latents. VGT-AE is a hybrid codec — a fine-tuned Qwen2.5-VL ViT as the encoder, a DC-AE decoder at sampling time. The tensor shape coincides with DC-AE f32c32 (32, 16, 16), but the space is different.… See the full description on the dataset page: https://huggingface.co/datasets/dingshizhe/vgtae-minit2i-sft-latents.
019
