CoolFace
Datasetpublic

augustinian-babylm/token-embeddings

token-embeddings Per-token visual embedding tables for the Augustinian BabyLM project: [V, 768] float32 matrices used to initialize the input embedding matrix of a DeBERTa-v3-base masked LM before text training. Organized as <encoder>/<vocab>/, for encoder in dinov3 / sam / ibot and vocab in 50k / 75k / 100k. Each directory holds E_init.safetensors (the table) and a seeded_mask marking which rows carry visual information, roughly 24-38% of rows depending on vocabulary size.… See the full description on the dataset page: https://huggingface.co/datasets/augustinian-babylm/token-embeddings.

sourceHugging Facecc-by-4.0updated 16d agoView on Hugging Face
0likes196downloads
14 commits on main
87cbc5416d ago

Add card

bylin
2360c472mo ago

Upload folder using huggingface_hub

bylin
c0666412mo ago

Upload folder using huggingface_hub

bylin
6bef6cd2mo ago

Upload folder using huggingface_hub

bylin
3e59f4c3mo ago

Upload folder using huggingface_hub

bylin
81085a33mo ago

Upload folder using huggingface_hub

bylin
64ee27b3mo ago

Upload folder using huggingface_hub

bylin
a68036b3mo ago

Upload folder using huggingface_hub

bylin
503ee9b3mo ago

Upload folder using huggingface_hub

bylin
ba51d263mo ago

Upload folder using huggingface_hub

bylin
5476c693mo ago

Upload folder using huggingface_hub

bylin
47787f13mo ago

Upload folder using huggingface_hub

bylin
647d0f13mo ago

Upload folder using huggingface_hub

bylin
3d3795b3mo ago

initial commit

bylin