datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tashkeela
Dataset Card for Tashkeela
Dataset Summary
It contains 75 million of fully vocalized words mainly
97 books from classical and modern Arabic language.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
The dataset is based on Arabic.
Dataset Structure
Data Instances
{'book':… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/tashkeela.arabic-msa-25k-saudi-male-tashkeel
Arabic MSA 25K — Saudi Male (Tashkeel)
25,000 fully-diacritized Arabic MSA text + audio pairs, rendered with a single
Saudi male neural voice at 48 kHz / 16-bit PCM, across 10 thematic categories.
Dataset Summary
arabic-msa-25k-saudi-male-tashkeel is a 25,000-clip Modern Standard Arabic (MSA)
speech corpus with matching diacritized text (full tashkeel / ḥarakāt). Every clip
is synthesized by the single voice ar-SA-HamedNeural (Azure Neural TTS, Saudi
Arabic male) at 48… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-msa-25k-saudi-male-tashkeel.tashkeel_deduparabic-tashkeel-dataset
Arabic Tashkeel Dataset
This is a fairly large dataset gathered from five main sources:
tashkeela (1.79GB - 45.05%): The entire Tashkeela dataset, repurposed in sentences. Some rows were omitted as they contain low diacritic (tashkeel characters) rate.
shamela (1.67GB - 42.10%): Random pages from over 2,000 books on the Shamela Library. Pages were selected using the below function (high diacritics rate)
wikipedia (269.94MB - 6.64%): A collection of Wikipedia articles. Diacritics… See the full description on the dataset page: https://huggingface.co/datasets/Abdou/arabic-tashkeel-dataset.Sadeed_Tashkeela
📚 Sadeed Tashkeela Arabic Diacritization Dataset
The Sadeed dataset is a large, high-quality Arabic diacritized corpus optimized for training and evaluating Arabic diacritization models.It is built exclusively from the Tashkeela corpus for the training set and a refined version of the Fadel Tashkeela test set for the test set.
Dataset Overview
Training Data:
Source: Cleaned version of the Tashkeela corpus (original data is ~75 million words, mostly Classical… See the full description on the dataset page: https://huggingface.co/datasets/Misraj/Sadeed_Tashkeela.tashkeela
Dataset Card for "tashkeela"
More Information needed
Arabic-Tashkeel-forkarabic-tashkeel-speech
Nahw Arabic Tashkeel Speech Dataset
An open-source collection of 1,093 fully diacritized Arabic speech recordings, crowd-sourced from native speakers via Nahw.ai.
Dataset summary
Stat
Value
Total recordings
1,093
Speakers
10
Language
Arabic (ar)
Sampling rate
16 kHz
License
CC-BY-4.0
Features
audio: The speech recording, resampled to 16 kHz.
transcription: The fully diacritized Arabic sentence that was read aloud.
sentence: The same… See the full description on the dataset page: https://huggingface.co/datasets/NahwAI/arabic-tashkeel-speech.Arabic-TashkeelFully diacritized arabic sentences gathered and cleaned from:
1st Dataset
and 2nd Dataset
where the Tashkeela part was removed from the first dataset and replaced with that of the second one.
tashkeelav2
Dataset Card for "tashkeelav2"
More Information needed
Tashkeela
Dataset Card for "Tashkeela"
More Information needed
Tashkeela
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Content
A version of the Tashkeela Arabic diacritized text dataset cleaned from the non-Arabic content and the undiacritized text, then divided into training, development, and testing sets.
The cleaning process includes removing the XML tags and strange symbols, as well as fixing… See the full description on the dataset page: https://huggingface.co/datasets/EmanKhater/Tashkeela.Fine-Tashkeel-CATT_benchmarkDIA2-Tashkeela-Gold
DIA2: A Comprehensive and Diverse Diacritized Modern Standard Arabic Corpus
DIA2 is a large-scale, natively sourced, and diacritized Modern Standard Arabic corpus
designed for NLP research and LLM development. It is curated from 28 diverse Arabic
sources including books, news articles, encyclopedic content, and poetry, and explicitly
avoids machine-translated content.
This repository contains the Tashkeela-Gold subset of DIA2.
The full DIA2 release consists of three datasets:… See the full description on the dataset page: https://huggingface.co/datasets/DIA2-Arabic/DIA2-Tashkeela-Gold.tashkeelashaar-tashkeel
Ashaar — Model-Inferred Tashkeel
Model-inferred diacritization (tashkeel) of 6,934,210 Arabic poetry verses
from the arbml/ashaar dataset.
Source
Raw verses: arbml/ashaar (241,964 poems, ~6.98M verse lines).
Diacritization model: basharalrfooh/Fine-Tashkeel
(ByT5-Large, fine-tuned on classical Arabic Tashkeela corpus).
Inference settings: FP16 on a single NVIDIA RTX 5880 Ada (48 GB),
max_new_tokens=48, max_input_len=128, greedy decoding. ~19 hours end-to-end
at ~100… See the full description on the dataset page: https://huggingface.co/datasets/mysamai/ashaar-tashkeel.tashkeel-arabic-sentencesThis dataset contains Arabic sentences extracted from the ImruQays/Alukah-Arabic dataset.
Sentences were filtered based on their 'tashkeel' (Arabic diacritics) ratio,
with a minimum ratio of 0.3 (adjustable during extraction).
Source:
The original articles were sourced from the ImruQays/Alukah-Arabic dataset on Hugging Face.
Processing:
Articles were loaded from ImruQays/Alukah-Arabic.
Each article was split into individual sentences using a regex pattern.
For each sentence, the ratio of… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/tashkeel-arabic-sentences.SadeedDiac-25_predictions_Fine-Tashkeelal_sallom_UAE_transcription_by_elevenlab_tashkeeltashkeel-dialect-pairs-v0tashkeelaFine-Tashkeel-SadeedDiac-25-predictionsCATT_benchmark_predictions_Fine-Tashkeeltashkeel
Arabic Tashkeel Dataset — Al-Maktaba Al-Shamela
A large-scale Arabic diacritization (tashkeel) dataset derived from Al-Maktaba Al-Shamela (المكتبة الشاملة), a comprehensive digital library of classical Islamic texts. The dataset pairs undiacritized Arabic sentences with their fully diacritized equivalents, enabling training and evaluation of automatic tashkeel systems.
Dataset Summary
Split
Examples
train
3,183,238
validation
397,904
test
397,906… See the full description on the dataset page: https://huggingface.co/datasets/yuujiElfahkrany/tashkeel.roots_ar_tashkeelaROOTS Subset: roots_ar_tashkeela
Tashkeela
Dataset uid: tashkeela
Description
The dataset collected from 97 books in both modern and classic arabic. The dataset contains Arabic diacritics. The dataset is
Homepage
https://sourceforge.net/projects/tashkeela/
Licensing
gpl-2.0: GNU General Public License v2.0 only
Speaker Locations
Sizes
0.2533 % of total
2.3340 % of ar
BigScience processing steps
Filters… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_ar_tashkeela.Tashkeel_corpustashkeela-audio
