augustinian-babylm/token-embeddings
token-embeddings Per-token visual embedding tables for the Augustinian BabyLM project: [V, 768] float32 matrices used to initialize the input embedding matrix of a DeBERTa-v3-base masked LM before text training. Organized as <encoder>/<vocab>/, for encoder in dinov3 / sam / ibot and vocab in 50k / 75k / 100k. Each directory holds E_init.safetensors (the table) and a seeded_mask marking which rows carry visual information, roughly 24-38% of rows depending on vocabulary size.… See the full description on the dataset page: https://huggingface.co/datasets/augustinian-babylm/token-embeddings.
0196
