CoolFace
Datasetpublic

dingshizhe/molmo-ae-cc12m-latents

CC12M latents — VGT-AE with a frozen SigLIP2 (Molmo) encoder Pre-encoded CC12M images as (32, 16, 16) float16 latents, paired with their captions. WebDataset format: 1097 tars, each member a <key>.npy + <key>.txt pair. samples 6,836,022 shards 1097 latent shape (32, 16, 16) float16 input resolution 512 px encoder VGTAE_Siglip2, siglip2vit_frozen_stage2 (50K steps) ViT Molmo SigLIP2-base/patch16, frozen through both stages decoder (for reference)… See the full description on the dataset page: https://huggingface.co/datasets/dingshizhe/molmo-ae-cc12m-latents.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
1likes659downloads
9 commits on main
5b3ad8c2mo ago

Add files using upload-large-folder tool

dingshizhe
fd1e0b62mo ago

Add files using upload-large-folder tool

dingshizhe
6a729bb2mo ago

Add files using upload-large-folder tool

dingshizhe
50fcf5d2mo ago

Add files using upload-large-folder tool

dingshizhe
5132edb2mo ago

Add files using upload-large-folder tool

dingshizhe
d06dde92mo ago

Add files using upload-large-folder tool

dingshizhe
03ae27a2mo ago

Add files using upload-large-folder tool

dingshizhe
3f73cee2mo ago

Add files using upload-large-folder tool

dingshizhe
40a30242mo ago

initial commit

dingshizhe