dingshizhe/molmo-ae-cc12m-latents
CC12M latents — VGT-AE with a frozen SigLIP2 (Molmo) encoder Pre-encoded CC12M images as (32, 16, 16) float16 latents, paired with their captions. WebDataset format: 1097 tars, each member a <key>.npy + <key>.txt pair. samples 6,836,022 shards 1097 latent shape (32, 16, 16) float16 input resolution 512 px encoder VGTAE_Siglip2, siglip2vit_frozen_stage2 (50K steps) ViT Molmo SigLIP2-base/patch16, frozen through both stages decoder (for reference)… See the full description on the dataset page: https://huggingface.co/datasets/dingshizhe/molmo-ae-cc12m-latents.
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
initial commit
