nilq/babylm-100M
BabyLM 100M This curated dataset is originally from the BabyLM Challenge. It consists of ~100M words of mixed domain, consisting of the following sources: CHILDES (child-directed speech) Subtitles (speech) BNC (speech) TED talks (speech) children's books (simple written language)
0112
BabyLM 100M
This curated dataset is originally from the BabyLM Challenge.
It consists of ~100M words of mixed domain, consisting of the following sources:
- CHILDES (child-directed speech)
- Subtitles (speech)
- BNC (speech)
- TED talks (speech)
- children's books (simple written language)
