CoolFace
Datasetpublic

adzcai/babylm-eng-nld-50-50-stratified

BabyLM English–Dutch 50/50 Stratified A bilingual training corpus for the Multilingual BabyLM project, combining 50% of the English and Dutch BabyBabelLM corpora. Sampling is stratified by category: each source category is sampled independently to preserve the original category proportions within each language. Token counts Language Tokens Share English (eng) 49,481,353 47.4% Dutch (nld) 54,953,487 52.6% Total 104,434,840 100%… See the full description on the dataset page: https://huggingface.co/datasets/adzcai/babylm-eng-nld-50-50-stratified.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes9downloads

adzcai/babylm-eng-nld-50-50-stratified · main · files are served by the source, never re-hosted here