miguelcsx/babylm-2026-mosaic-bpe-corpus
MOSAIC BBPE16K Encoded Corpus Reproducible encoded training artifact shared by the MOSAIC D384 model family. It contains the fixed 10M-word Strict-Small corpus representation used for a controlled 100M-word exposure schedule. Contents packed token and word arrays for training and validation; document offsets and token frequencies; document-complexity metadata; lexical recombination neighbors used by the variation-set arms; manifest.json with source and artifact… See the full description on the dataset page: https://huggingface.co/datasets/miguelcsx/babylm-2026-mosaic-bpe-corpus.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face