miguelcsx/babylm-2026-mosaic-bpe-corpus
MOSAIC BBPE16K Encoded Corpus Reproducible encoded training artifact shared by the MOSAIC D384 model family. It contains the fixed 10M-word Strict-Small corpus representation used for a controlled 100M-word exposure schedule. Contents packed token and word arrays for training and validation; document offsets and token frequencies; document-complexity metadata; lexical recombination neighbors used by the variation-set arms; manifest.json with source and artifact… See the full description on the dataset page: https://huggingface.co/datasets/miguelcsx/babylm-2026-mosaic-bpe-corpus.
016
