CoolFace
Datasetpublic

nilq/babylm-100M

BabyLM 100M This curated dataset is originally from the BabyLM Challenge. It consists of ~100M words of mixed domain, consisting of the following sources: CHILDES (child-directed speech) Subtitles (speech) BNC (speech) TED talks (speech) children's books (simple written language)

sourceHugging Faceupdated 3y agoView on Hugging Face
0likes112downloads

nilq/babylm-100M · main · files are served by the source, never re-hosted here