CoolFace
Datasetpublic

nilq/babylm-100M

BabyLM 100M This curated dataset is originally from the BabyLM Challenge. It consists of ~100M words of mixed domain, consisting of the following sources: CHILDES (child-directed speech) Subtitles (speech) BNC (speech) TED talks (speech) children's books (simple written language)

sourceHugging Faceupdated 3y agoView on Hugging Face
0likes112downloads
settings

This repository belongs to nilq on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namebabylm-100M
visibilitypublic
licencenot set
gatedno
ownernilq
Account settings
nilq/babylm-100M · CoolFace