datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
shamela_books_text_full
Shamela_Books_Text_Full
This dataset contains the full text content of Islamic Arabic books from the Shamela Library, organized by category, book, volume, and page, with footnotes stored separately. It is designed to support Arabic NLP, digital humanities, and bibliographic analysis.
🔗 This dataset is linked to the companion metadata dataset:
👉 Shamela_Books_info via the book_id field.
Update :
The dataset includes the original raw files as well as a single… See the full description on the dataset page: https://huggingface.co/datasets/MoMonir/shamela_books_text_full.Shamela_Books_info
Shamela Books information
This dataset contains structured metadata for 8,492 books sourced from the Shamela Library, with enhancements for clarity, consistency, and usability. It is intended to support NLP, bibliographic research, and digital humanities efforts involving Arabic texts.For full books text dataset please check shamela_books_text
Dataset Features
The dataset includes the following cleaned and standardized features:
Unification of Author Names: Author… See the full description on the dataset page: https://huggingface.co/datasets/MoMonir/Shamela_Books_info.shamela_books_text
Shamela_Books_Text
This dataset contains the full text content of Islamic Arabic books from the Shamela Library, organized by category, book, volume, and page, with footnotes stored separately. It is designed to support Arabic NLP, digital humanities, and bibliographic analysis.
🔗 This dataset is linked to the companion metadata dataset:
👉 Shamela_Books_info via the book_id field.
📊 Dataset Summary
Total Categories: 40
Total Books: 8,538
Total Records (Pages): 7… See the full description on the dataset page: https://huggingface.co/datasets/MoMonir/shamela_books_text.shamela_books_text_full
Shamela_Books_Text_Full
This dataset contains the full text content of Islamic Arabic books from the Shamela Library, organized by category, book, volume, and page, with footnotes stored separately. It is designed to support Arabic NLP, digital humanities, and bibliographic analysis.
🔗 This dataset is linked to the companion metadata dataset:
👉 Shamela_Books_info via the book_id field.
Update :
The dataset includes the original raw files as well as a single merged… See the full description on the dataset page: https://huggingface.co/datasets/mhaamh19/shamela_books_text_full.
