datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
token-embeddings
token-embeddings
Per-token visual embedding tables for the Augustinian BabyLM project: [V, 768]
float32 matrices used to initialize the input embedding matrix of a DeBERTa-v3-base
masked LM before text training.
Organized as <encoder>/<vocab>/, for encoder in dinov3 / sam / ibot and vocab in
50k / 75k / 100k. Each directory holds E_init.safetensors (the table) and a
seeded_mask marking which rows carry visual information, roughly 24-38% of rows
depending on vocabulary size.… See the full description on the dataset page: https://huggingface.co/datasets/augustinian-babylm/token-embeddings.babylm-2026-fr-92m-5seed-resultsbabylm-opensub-bilingual-50M
babylm-opensub-bilingual-50M
Bilingual OpenSubtitles training corpora (50M EN anchor, byte-premium partners).
Interleaved sentence streams for aligned/offset mixing modes.
Configs
en_nld_aligned
pair: en-nl / mode: aligned
pairs selected: 7303811
anchor units: 55796159
mixed sentences: 14607622
en_nld_offset
pair: en-nl / mode: offset
pairs selected: 7303811
anchor units: 55796159
mixed sentences: 14570422… See the full description on the dataset page: https://huggingface.co/datasets/zhzh98/babylm-opensub-bilingual-50M.vpswap-checkpoint-scores
VP-Swap checkpoint scores
Per-item correctness on the VP-Swap benchmark for nine models at twenty
points in training. This is the raw material behind Figures 4-6 of
Augustinian BabyLM (paper, code).
Layout
<model>/<revision>.jsonl, one line per benchmark item:
{"property": "color", "line": 6, "which": 1,
"pll_orig": -21.4213, "pll_swap": -23.9077, "correct": true}
property and line identify the source line in
eval/vpswap_bb24/vp_swap_<property>_pairs.txt; which… See the full description on the dataset page: https://huggingface.co/datasets/augustinian-babylm/vpswap-checkpoint-scores.dataset-biased-540Kdataset-biased-1.4Mdataset-anomalies-5.5Kdataset-biased-nvfp4-18Kdataset-biased-1.4M-pairwise-similaritybabylm-binomials-alldataset-biased-890Kbabylm_complexity_metricsdataset-anomalies-13Kdataset-qa-20Kdataset-qa-9Kbabylm-10m-easy-medium-hard-sortedbabylm-binomials-strictBabyLM_Dataset_291025babylm-tagged-by-common-complexity-metrics
