tashkeel
Datasets
All datasets matching “tashkeel”tashkeela
Dataset Card for Tashkeela
Dataset Summary
It contains 75 million of fully vocalized words mainly
97 books from classical and modern Arabic language.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
The dataset is based on Arabic.
Dataset Structure
Data Instances
{'book':… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/tashkeela.arabic-msa-25k-saudi-male-tashkeel
Arabic MSA 25K — Saudi Male (Tashkeel)
25,000 fully-diacritized Arabic MSA text + audio pairs, rendered with a single
Saudi male neural voice at 48 kHz / 16-bit PCM, across 10 thematic categories.
Dataset Summary
arabic-msa-25k-saudi-male-tashkeel is a 25,000-clip Modern Standard Arabic (MSA)
speech corpus with matching diacritized text (full tashkeel / ḥarakāt). Every clip
is synthesized by the single voice ar-SA-HamedNeural (Azure Neural TTS, Saudi
Arabic male) at 48… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-msa-25k-saudi-male-tashkeel.tashkeel_deduparabic-tashkeel-dataset
Arabic Tashkeel Dataset
This is a fairly large dataset gathered from five main sources:
tashkeela (1.79GB - 45.05%): The entire Tashkeela dataset, repurposed in sentences. Some rows were omitted as they contain low diacritic (tashkeel characters) rate.
shamela (1.67GB - 42.10%): Random pages from over 2,000 books on the Shamela Library. Pages were selected using the below function (high diacritics rate)
wikipedia (269.94MB - 6.64%): A collection of Wikipedia articles. Diacritics… See the full description on the dataset page: https://huggingface.co/datasets/Abdou/arabic-tashkeel-dataset.Sadeed_Tashkeela
📚 Sadeed Tashkeela Arabic Diacritization Dataset
The Sadeed dataset is a large, high-quality Arabic diacritized corpus optimized for training and evaluating Arabic diacritization models.It is built exclusively from the Tashkeela corpus for the training set and a refined version of the Fadel Tashkeela test set for the test set.
Dataset Overview
Training Data:
Source: Cleaned version of the Tashkeela corpus (original data is ~75 million words, mostly Classical… See the full description on the dataset page: https://huggingface.co/datasets/Misraj/Sadeed_Tashkeela.tashkeela
Dataset Card for "tashkeela"
More Information needed
