yuujiElfahkrany/tashkeel
Arabic Tashkeel Dataset — Al-Maktaba Al-Shamela A large-scale Arabic diacritization (tashkeel) dataset derived from Al-Maktaba Al-Shamela (المكتبة الشاملة), a comprehensive digital library of classical Islamic texts. The dataset pairs undiacritized Arabic sentences with their fully diacritized equivalents, enabling training and evaluation of automatic tashkeel systems. Dataset Summary Split Examples train 3,183,238 validation 397,904 test 397,906… See the full description on the dataset page: https://huggingface.co/datasets/yuujiElfahkrany/tashkeel.
This repository belongs to yuujiElfahkrany on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
