ebooks
Datasets
All datasets matching “ebooks”awesome-markdown-ebooks
Awesome-markdown-ebooks
Your GitHub PDFs, Now AI-Ready.
Project repo: https://github.com/OpenDataLab/awesome-markdown-ebooks
ara-ebooks-heavy-8.3M
BabyLM Arabic - Modified Corpora
Five alternative ~8.3M-word compositions of the BabyLM Arabic corpus
(BabyLM-community/babylm-ara), built to address that dataset's over-reliance on
OpenSubtitles (42%) and Habibi songs (37%) - together almost 80% of the original corpus -
by substituting in Arabic e-book content instead. See
jumelet2025babybabellm for the original dataset
and the wider BabyBabelLM project.
All five keep a fixed ~20% base of the original corpus's most… See the full description on the dataset page: https://huggingface.co/datasets/kabaros/ara-ebooks-heavy-8.3M.ara-ebooks-movies5-8.3M
BabyLM Arabic - Modified Corpora
Five alternative ~8.3M-word compositions of the BabyLM Arabic corpus
(BabyLM-community/babylm-ara), built to address that dataset's over-reliance on
OpenSubtitles (42%) and Habibi songs (37%) - together almost 80% of the original corpus -
by substituting in Arabic e-book content instead. See
jumelet2025babybabellm for the original dataset
and the wider BabyBabelLM project.
All five keep a fixed ~20% base of the original corpus's most… See the full description on the dataset page: https://huggingface.co/datasets/kabaros/ara-ebooks-movies5-8.3M.ara-ebooks-habibi5-8.3M
BabyLM Arabic - Modified Corpora
Five alternative ~8.3M-word compositions of the BabyLM Arabic corpus
(BabyLM-community/babylm-ara), built to address that dataset's over-reliance on
OpenSubtitles (42%) and Habibi songs (37%) - together almost 80% of the original corpus -
by substituting in Arabic e-book content instead. See
jumelet2025babybabellm for the original dataset
and the wider BabyBabelLM project.
All five keep a fixed ~20% base of the original corpus's most… See the full description on the dataset page: https://huggingface.co/datasets/kabaros/ara-ebooks-habibi5-8.3M.ebooksebooks
