CoolFace
Datasetpublicgated

Misraj/Sadeed_Tashkeela

📚 Sadeed Tashkeela Arabic Diacritization Dataset The Sadeed dataset is a large, high-quality Arabic diacritized corpus optimized for training and evaluating Arabic diacritization models.It is built exclusively from the Tashkeela corpus for the training set and a refined version of the Fadel Tashkeela test set for the test set. Dataset Overview Training Data: Source: Cleaned version of the Tashkeela corpus (original data is ~75 million words, mostly Classical… See the full description on the dataset page: https://huggingface.co/datasets/Misraj/Sadeed_Tashkeela.

sourceHugging Faceupdated 7d agoView on Hugging Face
16likes146downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.