CoolFace
Datasetpublic

miguelcsx/babylm-2026-mosaic-bpe-corpus

MOSAIC BBPE16K Encoded Corpus Reproducible encoded training artifact shared by the MOSAIC D384 model family. It contains the fixed 10M-word Strict-Small corpus representation used for a controlled 100M-word exposure schedule. Contents packed token and word arrays for training and validation; document offsets and token frequencies; document-complexity metadata; lexical recombination neighbors used by the variation-set arms; manifest.json with source and artifact… See the full description on the dataset page: https://huggingface.co/datasets/miguelcsx/babylm-2026-mosaic-bpe-corpus.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes16downloads
4 commits on main
65d36362mo ago

docs: update MOSAIC model links

miguelcsx
0334f612mo ago

docs: add MOSAIC corpus card

miguelcsx
c4ccb763mo ago

MOSAIC BPE-16384 corpus

miguelcsx
3e1aeab3mo ago

initial commit

miguelcsx