CoolFace
20 results

vgt

dingshizhe /vgtae-cc12m-latents CC12M latents encoded with VGT-AE (448px) Pre-encoded CC12M images in the VGT-AE latent space, so text-to-image training can skip the encoder entirely. These are not DC-AE latents. VGT-AE is a hybrid codec — a fine-tuned Qwen2.5-VL ViT as the encoder, a DC-AE decoder at sampling time. The tensor shape happens to match DC-AE f32c32 (32, 16, 16), but the space is completely different; mixing the two silently trains a model against noise. Contents 1097 WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/dingshizhe/vgtae-cc12m-latents.text-to-image1M<n<10M0 likes651 downloads1mo agoHugging Facesionic-ai /VGT-Kor-FineSocial-QA0 likes16 downloads9mo agoHugging Facedingshizhe /vgtae-minit2i-sft-latents MiniT2I SFT-set latents encoded with VGT-AE (448px) The text-to-image SFT mixture, pre-encoded into the VGT-AE latent space. Companion to dingshizhe/vgtae-cc12m-latents, which covers the pretraining corpus; same encoder, same settings, same layout. These are not DC-AE latents. VGT-AE is a hybrid codec — a fine-tuned Qwen2.5-VL ViT as the encoder, a DC-AE decoder at sampling time. The tensor shape coincides with DC-AE f32c32 (32, 16, 16), but the space is different.… See the full description on the dataset page: https://huggingface.co/datasets/dingshizhe/vgtae-minit2i-sft-latents.text-to-image10K<n<100K0 likes8 downloads1mo agoHugging Facetaufik89 /vgt5t50 likes4 downloads8mo agoHugging Facedaaxila /twitter-jaNCb7zmGFgi9Um-2023.11.07-1721955848426275168-vGt-D8RXdloL7ueM-part1imagen<1K0 likes4 downloads5mo agoHugging Facethanhtungd /vgtre0 likes2 downloads9mo agoHugging Face