CoolFace
20 results

tashkeel

community-datasets /tashkeela Dataset Card for Tashkeela Dataset Summary It contains 75 million of fully vocalized words mainly 97 books from classical and modern Arabic language. Supported Tasks and Leaderboards [More Information Needed] Languages The dataset is based on Arabic. Dataset Structure Data Instances {'book':… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/tashkeela.texttext-generationn<1K6 likes222 downloads2y agoHugging FaceHeshamHaroon /arabic-msa-25k-saudi-male-tashkeel Arabic MSA 25K — Saudi Male (Tashkeel) 25,000 fully-diacritized Arabic MSA text + audio pairs, rendered with a single Saudi male neural voice at 48 kHz / 16-bit PCM, across 10 thematic categories. Dataset Summary arabic-msa-25k-saudi-male-tashkeel is a 25,000-clip Modern Standard Arabic (MSA) speech corpus with matching diacritized text (full tashkeel / ḥarakāt). Every clip is synthesized by the single voice ar-SA-HamedNeural (Azure Neural TTS, Saudi Arabic male) at 48… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-msa-25k-saudi-male-tashkeel.tabulartext-to-speech10K<n<100K10 likes186 downloads5mo agoHugging Faceohsn /tashkeel_deduptext1M<n<10M0 likes179 downloads7mo agoHugging FaceAbdou /arabic-tashkeel-dataset Arabic Tashkeel Dataset This is a fairly large dataset gathered from five main sources: tashkeela (1.79GB - 45.05%): The entire Tashkeela dataset, repurposed in sentences. Some rows were omitted as they contain low diacritic (tashkeel characters) rate. shamela (1.67GB - 42.10%): Random pages from over 2,000 books on the Shamela Library. Pages were selected using the below function (high diacritics rate) wikipedia (269.94MB - 6.64%): A collection of Wikipedia articles. Diacritics… See the full description on the dataset page: https://huggingface.co/datasets/Abdou/arabic-tashkeel-dataset.text1M<n<10M7 likes146 downloads2y agoHugging FaceMisraj /Sadeed_Tashkeelagated 📚 Sadeed Tashkeela Arabic Diacritization Dataset The Sadeed dataset is a large, high-quality Arabic diacritized corpus optimized for training and evaluating Arabic diacritization models.It is built exclusively from the Tashkeela corpus for the training set and a refined version of the Fadel Tashkeela test set for the test set. Dataset Overview Training Data: Source: Cleaned version of the Tashkeela corpus (original data is ~75 million words, mostly Classical… See the full description on the dataset page: https://huggingface.co/datasets/Misraj/Sadeed_Tashkeela.texttext-generation1M<n<10M16 likes146 downloads7d agoHugging Facearbml /tashkeela Dataset Card for "tashkeela" More Information needed text1M<n<10M7 likes130 downloads3y agoHugging Face