MoMonir/shamela_books_text_full
Shamela_Books_Text_Full This dataset contains the full text content of Islamic Arabic books from the Shamela Library, organized by category, book, volume, and page, with footnotes stored separately. It is designed to support Arabic NLP, digital humanities, and bibliographic analysis. 🔗 This dataset is linked to the companion metadata dataset: 👉 Shamela_Books_info via the book_id field. Update : The dataset includes the original raw files as well as a single… See the full description on the dataset page: https://huggingface.co/datasets/MoMonir/shamela_books_text_full.
Upload dataset files from category 001 to 040
chore: remove old sequential shards before category reorganization
Update README.md
Update README to support merged and sharded configs
Upload merged parquet file: shamela_full_merged.parquet
Delete shamela_books_count.png
Upload shamela_books_count.png with huggingface_hub
Update README.md
Update README.md
Delete shamela_books_text_full.csv
Delete dataset_infos.json
Upload dataset (part 00001-of-00002)
Upload dataset (part 00000-of-00002)
Upload dataset
Update README.md
Update README.md
Create README.md
Update dataset_infos.json
Create dataset_infos.json
Delete dataset_info.yaml
Delete .hf
Create .hf
Upload dataset_info.yaml
Upload shamela_books_text_full.csv with huggingface_hub
initial commit
