CoolFace
Datasetpublic

cambridge-climb/BabyLM

Dataset for the shared baby language modeling task. The goal is to train a language model from scratch on this data which represents roughly the amount of text and speech data a young child observes.

sourceHugging Faceupdated 2y agoView on Hugging Face
3likes2.8kdownloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
cambridge-climb/BabyLM · CoolFace