daichi812/babyloop-mm-corpus
Datasheet: babyloop Multimodal Training Corpus (BabyLM 2026, Strict track) Datasheet format follows Gebru et al. (2021), "Datasheets for Datasets," abridged to the sections relevant for a training-corpus release. All counts below are measured, not estimated; provenance manifests are versioned in this repository under manifests/. Summary Total word budget 99,999,984 whitespace words (≤ 100M, BabyLM Strict rule) Text portion 49,999,988 words —… See the full description on the dataset page: https://huggingface.co/datasets/daichi812/babyloop-mm-corpus.
Update code repository links to the new GitHub username (daichi0812 -> daichi8120)
Link the public assembly code repository (github.com/daichi0812/babyloop)
Fix Loading snippet: captions config is served by the text builder; parse JSON lines explicitly
Add datasheet card, assembled text/caption corpus, and provenance manifests
initial commit
