CoolFace
Datasetpublic

kabaros/ara-ebooks-habibi5-8.3M

BabyLM Arabic - Modified Corpora Five alternative ~8.3M-word compositions of the BabyLM Arabic corpus (BabyLM-community/babylm-ara), built to address that dataset's over-reliance on OpenSubtitles (42%) and Habibi songs (37%) - together almost 80% of the original corpus - by substituting in Arabic e-book content instead. See jumelet2025babybabellm for the original dataset and the wider BabyBabelLM project. All five keep a fixed ~20% base of the original corpus's most… See the full description on the dataset page: https://huggingface.co/datasets/kabaros/ara-ebooks-habibi5-8.3M.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes18downloads
8 commits on main
c4f3afe1mo ago

Update README.md

kabaros
18c34eb1mo ago

Update README.md

kabaros
de838001mo ago

Upload babylm_results.png with huggingface_hub

kabaros
6fa0a361mo ago

Upload README.md with huggingface_hub

kabaros
326757b1mo ago

Upload README.md with huggingface_hub

kabaros
6bea35e1mo ago

Upload README.md with huggingface_hub

kabaros
252034a1mo ago

Upload dataset

kabaros
c6795e41mo ago

initial commit

kabaros