CoolFace
Datasetpublic

babylm-anon/stratified_10m_curriculum

Dataset Card for Stratified 10M Curriculum This is a stratified split by domain of the raw datasets used by the 2024 BabyLM challange. Sampled from the original datasets, equal amounts of tokens for each stage (C1,...C5). Child-directed speech accounts for nearly half of the original dataset by word count. In preliminary experiments using a training data influence estimation method, this category was by far the most influential. This dataset enables us to investigate whether… See the full description on the dataset page: https://huggingface.co/datasets/babylm-anon/stratified_10m_curriculum.

sourceHugging Facecc-by-sa-4.0updated 1y agoView on Hugging Face
0likes1.3kdownloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
babylm-anon/stratified_10m_curriculum · CoolFace