vgt
Datasets
All datasets matching “vgt”vgtae-cc12m-latents
CC12M latents encoded with VGT-AE (448px)
Pre-encoded CC12M
images in the VGT-AE latent space, so text-to-image training can skip the
encoder entirely.
These are not DC-AE latents. VGT-AE is a hybrid codec — a fine-tuned
Qwen2.5-VL ViT as the encoder, a DC-AE decoder at sampling time. The tensor shape
happens to match DC-AE f32c32 (32, 16, 16), but the space is completely
different; mixing the two silently trains a model against noise.
Contents
1097 WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/dingshizhe/vgtae-cc12m-latents.VGT-Kor-FineSocial-QAvgtae-minit2i-sft-latents
MiniT2I SFT-set latents encoded with VGT-AE (448px)
The text-to-image SFT mixture, pre-encoded into the VGT-AE latent space.
Companion to dingshizhe/vgtae-cc12m-latents,
which covers the pretraining corpus; same encoder, same settings, same layout.
These are not DC-AE latents. VGT-AE is a hybrid codec — a fine-tuned
Qwen2.5-VL ViT as the encoder, a DC-AE decoder at sampling time. The tensor
shape coincides with DC-AE f32c32 (32, 16, 16), but the space is different.… See the full description on the dataset page: https://huggingface.co/datasets/dingshizhe/vgtae-minit2i-sft-latents.vgt5t5twitter-jaNCb7zmGFgi9Um-2023.11.07-1721955848426275168-vGt-D8RXdloL7ueM-part1vgtre
