CoolFace
Datasetpublic

daichi812/babyloop-mm-corpus

Datasheet: babyloop Multimodal Training Corpus (BabyLM 2026, Strict track) Datasheet format follows Gebru et al. (2021), "Datasheets for Datasets," abridged to the sections relevant for a training-corpus release. All counts below are measured, not estimated; provenance manifests are versioned in this repository under manifests/. Summary Total word budget 99,999,984 whitespace words (≤ 100M, BabyLM Strict rule) Text portion 49,999,988 words —… See the full description on the dataset page: https://huggingface.co/datasets/daichi812/babyloop-mm-corpus.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes41downloads
5 commits on main
d17e99a1mo ago

Update code repository links to the new GitHub username (daichi0812 -> daichi8120)

daichi812
81054753mo ago

Link the public assembly code repository (github.com/daichi0812/babyloop)

daichi812
8f932b93mo ago

Fix Loading snippet: captions config is served by the text builder; parse JSON lines explicitly

daichi812
1d930433mo ago

Add datasheet card, assembled text/caption corpus, and provenance manifests

daichi812
e1e99a73mo ago

initial commit

daichi812