CoolFace
Datasetpublic

miguelcsx/babylm-2026-mosaic-bpe-corpus

MOSAIC BBPE16K Encoded Corpus Reproducible encoded training artifact shared by the MOSAIC D384 model family. It contains the fixed 10M-word Strict-Small corpus representation used for a controlled 100M-word exposure schedule. Contents packed token and word arrays for training and validation; document offsets and token frequencies; document-complexity metadata; lexical recombination neighbors used by the variation-set arms; manifest.json with source and artifact… See the full description on the dataset page: https://huggingface.co/datasets/miguelcsx/babylm-2026-mosaic-bpe-corpus.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes16downloads
Dataset Card

MOSAIC BBPE16K Encoded Corpus

Reproducible encoded training artifact shared by the MOSAIC D384 model family. It contains the fixed 10M-word Strict-Small corpus representation used for a controlled 100M-word exposure schedule.

Contents

  • —packed token and word arrays for training and validation;
  • —document offsets and token frequencies;
  • —document-complexity metadata;
  • —lexical recombination neighbors used by the variation-set arms;
  • —manifest.json with source and artifact provenance.

This repository contains encoded research artifacts, not an additional source of training text. The model repositories bundle the exact tokenizer needed for loading and evaluation.

Associated models