CoolFace
Datasetpublicgated

Misraj/Sadeed_Tashkeela

📚 Sadeed Tashkeela Arabic Diacritization Dataset The Sadeed dataset is a large, high-quality Arabic diacritized corpus optimized for training and evaluating Arabic diacritization models.It is built exclusively from the Tashkeela corpus for the training set and a refined version of the Fadel Tashkeela test set for the test set. Dataset Overview Training Data: Source: Cleaned version of the Tashkeela corpus (original data is ~75 million words, mostly Classical… See the full description on the dataset page: https://huggingface.co/datasets/Misraj/Sadeed_Tashkeela.

sourceHugging Faceupdated 10d agoView on Hugging Face
16likes169downloads
settings

This repository belongs to Misraj on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameSadeed_Tashkeela
visibilitypublic
licencenot set
gatedyes
ownerMisraj
Account settings
Misraj/Sadeed_Tashkeela · CoolFace