CoolFace
Datasetpublic

kabaros/ara-ebooks-habibi5-8.3M

BabyLM Arabic - Modified Corpora Five alternative ~8.3M-word compositions of the BabyLM Arabic corpus (BabyLM-community/babylm-ara), built to address that dataset's over-reliance on OpenSubtitles (42%) and Habibi songs (37%) - together almost 80% of the original corpus - by substituting in Arabic e-book content instead. See jumelet2025babybabellm for the original dataset and the wider BabyBabelLM project. All five keep a fixed ~20% base of the original corpus's most… See the full description on the dataset page: https://huggingface.co/datasets/kabaros/ara-ebooks-habibi5-8.3M.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes25downloads
Dataset Card

BabyLM Arabic - Modified Corpora

Five alternative ~8.3M-word compositions of the BabyLM Arabic corpus (BabyLM-community/babylm-ara), built to address that dataset's over-reliance on OpenSubtitles (42%) and Habibi songs (37%) - together almost 80% of the original corpus - by substituting in Arabic e-book content instead. See jumelet2025babybabellm for the original dataset and the wider BabyBabelLM project.

All five keep a fixed ~20% base of the original corpus's most developmentally plausible content - child-books and child-wiki - and vary only how the remaining ~80% is filled.

DatasetComposition
ebooks-heavy (custom-8.3M-wiki_0-movies_0-habibi_0-ebooks_100)Fixed content ~20%, e-books ~80%
ebooks w/ 5% habibi (custom-8.3M-wiki_0-movies_0-habibi_5-ebooks_95)Fixed content ~20%, e-books ~76%, Habibi songs ~4% (Egyptian dialect)
ebooks w/ 5% movies (custom-8.3M-wiki_0-movies_5-habibi_0-ebooks_95)Fixed content ~20%, e-books ~76%, movies ~4%
Impossible Man (custom-impossible-man-8.3M-fixed-20%-ebooks-42%-hindawi-38%)Fixed content 20.2%, e-books 42.2%, Hindawi YA fiction 37.7%
Diverse (diversity-8.3M-fixed-20%-movies-17%-habibi-17%-ebooks-8%-hindawi-38%)Fixed content 20.2%, movies 17%, Habibi songs (Egyptian) 17%, e-books 8%, Hindawi YA fiction 38%

"e-books" = the Arabic e-book corpus (Hallberg, 2024; Hindawi Foundation), extended with a broader subset of middle/high-school-appropriate categories (history, technology, arts, science, science fiction).

"movies" = a rebuilt OpenSubtitles pool sourced from OPUS OpenSubtitles, re-rated for child-suitability via IMDb ID lookup and filtered to ratings U through PG-13.

"Hindawi YA fiction" = The Man of the Impossible, a popular Arabic Young Adult book series from the Hindawi project not yet included in the main e-book corpus release.

All five were built and evaluated as part of an MSc dissertation on tokenisation and dataset composition for small-scale Arabic language models (University of Stirling). The morphologically-aware tokeniser/dataset combination trained on the ebooks-heavy variant achieved the best grammaticality (MultiBLiMP) scores among all variants tested.

[image]