CoolFace
Datasetpublic

nilq/babylm-100M

BabyLM 100M This curated dataset is originally from the BabyLM Challenge. It consists of ~100M words of mixed domain, consisting of the following sources: CHILDES (child-directed speech) Subtitles (speech) BNC (speech) TED talks (speech) children's books (simple written language)

sourceHugging Faceupdated 3y agoView on Hugging Face
0likes112downloads
Dataset Card

BabyLM 100M

This curated dataset is originally from the BabyLM Challenge.

It consists of ~100M words of mixed domain, consisting of the following sources:

  • —CHILDES (child-directed speech)
  • —Subtitles (speech)
  • —BNC (speech)
  • —TED talks (speech)
  • —children's books (simple written language)