CoolFace
Datasetpublic

adzcai/babylm-eng-nld-50-50-stratified

BabyLM English–Dutch 50/50 Stratified A bilingual training corpus for the Multilingual BabyLM project, combining 50% of the English and Dutch BabyBabelLM corpora. Sampling is stratified by category: each source category is sampled independently to preserve the original category proportions within each language. Token counts Language Tokens Share English (eng) 49,481,353 47.4% Dutch (nld) 54,953,487 52.6% Total 104,434,840 100%… See the full description on the dataset page: https://huggingface.co/datasets/adzcai/babylm-eng-nld-50-50-stratified.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes9downloads
Dataset Card

BabyLM English–Dutch 50/50 Stratified

A bilingual training corpus for the Multilingual BabyLM project, combining 50% of the English and Dutch BabyBabelLM corpora.

Sampling is stratified by `category`: each source category is sampled independently to preserve the original category proportions within each language.

Token counts

LanguageTokensShare
English (eng)49,481,35347.4%
Dutch (nld)54,953,48752.6%
Total104,434,840100%

Category breakdown

CategoryEnglish tokensEnglish %Dutch tokensDutch %
child-available-speech4,551,7679.2%00.0%
child-books13,406,59427.1%2,288,5234.2%
child-directed-speech13,857,31428.0%1,656,9983.0%
child-news00.0%1,589,0462.9%
child-wiki7,304,85614.8%4,836,7608.8%
educational00.0%9,573,10817.4%
padding-opensubtitles9,856,92219.9%34,523,37162.8%
padding-wikipedia503,9001.0%00.0%
subtitles00.0%485,6810.9%

Source datasets