CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kabaros /ara-ebooks-heavy-8.3M BabyLM Arabic - Modified Corpora Five alternative ~8.3M-word compositions of the BabyLM Arabic corpus (BabyLM-community/babylm-ara), built to address that dataset's over-reliance on OpenSubtitles (42%) and Habibi songs (37%) - together almost 80% of the original corpus - by substituting in Arabic e-book content instead. See jumelet2025babybabellm for the original dataset and the wider BabyBabelLM project. All five keep a fixed ~20% base of the original corpus's most… See the full description on the dataset page: https://huggingface.co/datasets/kabaros/ara-ebooks-heavy-8.3M.text1K<n<10K0 likes51 downloads1mo agoHugging Face02kabaros /ara-ebooks-movies5-8.3M BabyLM Arabic - Modified Corpora Five alternative ~8.3M-word compositions of the BabyLM Arabic corpus (BabyLM-community/babylm-ara), built to address that dataset's over-reliance on OpenSubtitles (42%) and Habibi songs (37%) - together almost 80% of the original corpus - by substituting in Arabic e-book content instead. See jumelet2025babybabellm for the original dataset and the wider BabyBabelLM project. All five keep a fixed ~20% base of the original corpus's most… See the full description on the dataset page: https://huggingface.co/datasets/kabaros/ara-ebooks-movies5-8.3M.text1K<n<10K0 likes39 downloads1mo agoHugging Face03kabaros /ara-ebooks-habibi5-8.3M BabyLM Arabic - Modified Corpora Five alternative ~8.3M-word compositions of the BabyLM Arabic corpus (BabyLM-community/babylm-ara), built to address that dataset's over-reliance on OpenSubtitles (42%) and Habibi songs (37%) - together almost 80% of the original corpus - by substituting in Arabic e-book content instead. See jumelet2025babybabellm for the original dataset and the wider BabyBabelLM project. All five keep a fixed ~20% base of the original corpus's most… See the full description on the dataset page: https://huggingface.co/datasets/kabaros/ara-ebooks-habibi5-8.3M.text10K<n<100K0 likes34 downloads1mo agoHugging Face04enelpol /gutenberg_selected_ebooks Gutenberg selected ebooks dataset This dataset is a collection of passages from ebooks handpicked from the Gutenberg Project. These writings are: Alice's Adventures in Wonderland Pride and Prejudice Romeo and Juliet The Adventures of Sherlock Holmes The Odyssey Winnie-the-Pooh Source The texts of the passages were derived from a larger Gutenberg-based set: sedthh/gutenberg_english, which was sourced directly from the project's site. Metadata Each passage… See the full description on the dataset page: https://huggingface.co/datasets/enelpol/gutenberg_selected_ebooks.tabularquestion-answering1K<n<10K1 likes13 downloads2y agoHugging Face05jakeveo05 /ebooks-multilingualtextn<1K0 likes7 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.