adzcai/babylm-eng-nld-50-50-stratified
BabyLM English–Dutch 50/50 Stratified A bilingual training corpus for the Multilingual BabyLM project, combining 50% of the English and Dutch BabyBabelLM corpora. Sampling is stratified by category: each source category is sampled independently to preserve the original category proportions within each language. Token counts Language Tokens Share English (eng) 49,481,353 47.4% Dutch (nld) 54,953,487 52.6% Total 104,434,840 100%… See the full description on the dataset page: https://huggingface.co/datasets/adzcai/babylm-eng-nld-50-50-stratified.
BabyLM English–Dutch 50/50 Stratified
A bilingual training corpus for the Multilingual BabyLM project, combining 50% of the English and Dutch BabyBabelLM corpora.
Sampling is stratified by `category`: each source category is sampled independently to preserve the original category proportions within each language.
