CoolFace
Datasetpublic

dingshizhe/vgtae-minit2i-sft-latents

MiniT2I SFT-set latents encoded with VGT-AE (448px) The text-to-image SFT mixture, pre-encoded into the VGT-AE latent space. Companion to dingshizhe/vgtae-cc12m-latents, which covers the pretraining corpus; same encoder, same settings, same layout. These are not DC-AE latents. VGT-AE is a hybrid codec — a fine-tuned Qwen2.5-VL ViT as the encoder, a DC-AE decoder at sampling time. The tensor shape coincides with DC-AE f32c32 (32, 16, 16), but the space is different.… See the full description on the dataset page: https://huggingface.co/datasets/dingshizhe/vgtae-minit2i-sft-latents.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes19downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
dingshizhe/vgtae-minit2i-sft-latents · CoolFace