CoolFace
Datasetpublic

miguelcsx/babylm-2026-mosaic-bpe-corpus

MOSAIC BBPE16K Encoded Corpus Reproducible encoded training artifact shared by the MOSAIC D384 model family. It contains the fixed 10M-word Strict-Small corpus representation used for a controlled 100M-word exposure schedule. Contents packed token and word arrays for training and validation; document offsets and token frequencies; document-complexity metadata; lexical recombination neighbors used by the variation-set arms; manifest.json with source and artifact… See the full description on the dataset page: https://huggingface.co/datasets/miguelcsx/babylm-2026-mosaic-bpe-corpus.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes16downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face