CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01community-datasets /tashkeela Dataset Card for Tashkeela Dataset Summary It contains 75 million of fully vocalized words mainly 97 books from classical and modern Arabic language. Supported Tasks and Leaderboards [More Information Needed] Languages The dataset is based on Arabic. Dataset Structure Data Instances {'book':… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/tashkeela.texttext-generationn<1K6 likes221 downloads2y agoHugging Face02HeshamHaroon /arabic-msa-25k-saudi-male-tashkeel Arabic MSA 25K — Saudi Male (Tashkeel) 25,000 fully-diacritized Arabic MSA text + audio pairs, rendered with a single Saudi male neural voice at 48 kHz / 16-bit PCM, across 10 thematic categories. Dataset Summary arabic-msa-25k-saudi-male-tashkeel is a 25,000-clip Modern Standard Arabic (MSA) speech corpus with matching diacritized text (full tashkeel / ḥarakāt). Every clip is synthesized by the single voice ar-SA-HamedNeural (Azure Neural TTS, Saudi Arabic male) at 48… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-msa-25k-saudi-male-tashkeel.tabulartext-to-speech10K<n<100K10 likes197 downloads5mo agoHugging Face03ohsn /tashkeel_deduptext1M<n<10M0 likes168 downloads7mo agoHugging Face04Abdou /arabic-tashkeel-dataset Arabic Tashkeel Dataset This is a fairly large dataset gathered from five main sources: tashkeela (1.79GB - 45.05%): The entire Tashkeela dataset, repurposed in sentences. Some rows were omitted as they contain low diacritic (tashkeel characters) rate. shamela (1.67GB - 42.10%): Random pages from over 2,000 books on the Shamela Library. Pages were selected using the below function (high diacritics rate) wikipedia (269.94MB - 6.64%): A collection of Wikipedia articles. Diacritics… See the full description on the dataset page: https://huggingface.co/datasets/Abdou/arabic-tashkeel-dataset.text1M<n<10M7 likes156 downloads2y agoHugging Face05Misraj /Sadeed_Tashkeelagated 📚 Sadeed Tashkeela Arabic Diacritization Dataset The Sadeed dataset is a large, high-quality Arabic diacritized corpus optimized for training and evaluating Arabic diacritization models.It is built exclusively from the Tashkeela corpus for the training set and a refined version of the Fadel Tashkeela test set for the test set. Dataset Overview Training Data: Source: Cleaned version of the Tashkeela corpus (original data is ~75 million words, mostly Classical… See the full description on the dataset page: https://huggingface.co/datasets/Misraj/Sadeed_Tashkeela.texttext-generation1M<n<10M16 likes153 downloads7d agoHugging Face06arbml /tashkeela Dataset Card for "tashkeela" More Information needed text1M<n<10M7 likes136 downloads3y agoHugging Face07Abdo-Alshoki /Arabic-Tashkeel-forktext1M<n<10M0 likes122 downloads1y agoHugging Face08NahwAI /arabic-tashkeel-speech Nahw Arabic Tashkeel Speech Dataset An open-source collection of 1,093 fully diacritized Arabic speech recordings, crowd-sourced from native speakers via Nahw.ai. Dataset summary Stat Value Total recordings 1,093 Speakers 10 Language Arabic (ar) Sampling rate 16 kHz License CC-BY-4.0 Features audio: The speech recording, resampled to 16 kHz. transcription: The fully diacritized Arabic sentence that was read aloud. sentence: The same… See the full description on the dataset page: https://huggingface.co/datasets/NahwAI/arabic-tashkeel-speech.audioautomatic-speech-recognition1K<n<10K3 likes69 downloads5mo agoHugging Face09CUAIStudents /Arabic-TashkeelFully diacritized arabic sentences gathered and cleaned from: 1st Dataset and 2nd Dataset where the Tashkeela part was removed from the first dataset and replaced with that of the second one. text1M<n<10M2 likes57 downloads1y agoHugging Face10arbml /tashkeelav2 Dataset Card for "tashkeelav2" More Information needed text100K<n<1M6 likes52 downloads3y agoHugging Face11asas-ai /Tashkeela Dataset Card for "Tashkeela" More Information needed text1M<n<10M1 likes42 downloads3y agoHugging Face12EmanKhater /Tashkeela Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Content A version of the Tashkeela Arabic diacritized text dataset cleaned from the non-Arabic content and the undiacritized text, then divided into training, development, and testing sets. The cleaning process includes removing the XML tags and strange symbols, as well as fixing… See the full description on the dataset page: https://huggingface.co/datasets/EmanKhater/Tashkeela.text1M<n<10M2 likes34 downloads2y agoHugging Face13MoaazTalab /Fine-Tashkeel-CATT_benchmarktextn<1K0 likes30 downloads1y agoHugging Face14DIA2-Arabic /DIA2-Tashkeela-Goldgated DIA2: A Comprehensive and Diverse Diacritized Modern Standard Arabic Corpus DIA2 is a large-scale, natively sourced, and diacritized Modern Standard Arabic corpus designed for NLP research and LLM development. It is curated from 28 diverse Arabic sources including books, news articles, encyclopedic content, and poetry, and explicitly avoids machine-translated content. This repository contains the Tashkeela-Gold subset of DIA2. The full DIA2 release consists of three datasets:… See the full description on the dataset page: https://huggingface.co/datasets/DIA2-Arabic/DIA2-Tashkeela-Gold.texttext-generation10K<n<100K0 likes27 downloads6mo agoHugging Face15whitefox123 /tashkeelaudio10K<n<100K3 likes24 downloads3y agoHugging Face16mysamai /ashaar-tashkeel Ashaar — Model-Inferred Tashkeel Model-inferred diacritization (tashkeel) of 6,934,210 Arabic poetry verses from the arbml/ashaar dataset. Source Raw verses: arbml/ashaar (241,964 poems, ~6.98M verse lines). Diacritization model: basharalrfooh/Fine-Tashkeel (ByT5-Large, fine-tuned on classical Arabic Tashkeela corpus). Inference settings: FP16 on a single NVIDIA RTX 5880 Ada (48 GB), max_new_tokens=48, max_input_len=128, greedy decoding. ~19 hours end-to-end at ~100… See the full description on the dataset page: https://huggingface.co/datasets/mysamai/ashaar-tashkeel.texttext-generation1M<n<10M0 likes22 downloads5mo agoHugging Face17riotu-lab /tashkeel-arabic-sentencesThis dataset contains Arabic sentences extracted from the ImruQays/Alukah-Arabic dataset. Sentences were filtered based on their 'tashkeel' (Arabic diacritics) ratio, with a minimum ratio of 0.3 (adjustable during extraction). Source: The original articles were sourced from the ImruQays/Alukah-Arabic dataset on Hugging Face. Processing: Articles were loaded from ImruQays/Alukah-Arabic. Each article was split into individual sentences using a regex pattern. For each sentence, the ratio of… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/tashkeel-arabic-sentences.texttranslation100K<n<1M0 likes21 downloads6mo agoHugging Face18Bisher /SadeedDiac-25_predictions_Fine-Tashkeeltext1K<n<10K0 likes17 downloads1y agoHugging Face19AhmedBadawy11 /al_sallom_UAE_transcription_by_elevenlab_tashkeelaudion<1K0 likes13 downloads1y agoHugging Face20mohammed-bahumaish /tashkeel-dialect-pairs-v0gatedtext10M<n<100M0 likes11 downloads7d agoHugging Face21glonor /tashkeelatext1M<n<10M1 likes9 downloads2y agoHugging Face22MoaazTalab /Fine-Tashkeel-SadeedDiac-25-predictionstext1K<n<10K0 likes9 downloads1y agoHugging Face23Bisher /CATT_benchmark_predictions_Fine-Tashkeeltextn<1K0 likes8 downloads1y agoHugging Face24yuujiElfahkrany /tashkeel Arabic Tashkeel Dataset — Al-Maktaba Al-Shamela A large-scale Arabic diacritization (tashkeel) dataset derived from Al-Maktaba Al-Shamela (المكتبة الشاملة), a comprehensive digital library of classical Islamic texts. The dataset pairs undiacritized Arabic sentences with their fully diacritized equivalents, enabling training and evaluation of automatic tashkeel systems. Dataset Summary Split Examples train 3,183,238 validation 397,904 test 397,906… See the full description on the dataset page: https://huggingface.co/datasets/yuujiElfahkrany/tashkeel.text1M<n<10M0 likes8 downloads3mo agoHugging Face25bigscience-data /roots_ar_tashkeelagatedROOTS Subset: roots_ar_tashkeela Tashkeela Dataset uid: tashkeela Description The dataset collected from 97 books in both modern and classic arabic. The dataset contains Arabic diacritics. The dataset is Homepage https://sourceforge.net/projects/tashkeela/ Licensing gpl-2.0: GNU General Public License v2.0 only Speaker Locations Sizes 0.2533 % of total 2.3340 % of ar BigScience processing steps Filters… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_ar_tashkeela.textn<1K0 likes6 downloads4y agoHugging Face26TashkeelCorpus /Tashkeel_corpustext1M<n<10M0 likes6 downloads7mo agoHugging Face27muhtasham /tashkeela-audiogatedaudio10K<n<100K1 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.