datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tashkeela
Dataset Card for Tashkeela
Dataset Summary
It contains 75 million of fully vocalized words mainly
97 books from classical and modern Arabic language.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
The dataset is based on Arabic.
Dataset Structure
Data Instances
{'book':… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/tashkeela.Sadeed_Tashkeela
📚 Sadeed Tashkeela Arabic Diacritization Dataset
The Sadeed dataset is a large, high-quality Arabic diacritized corpus optimized for training and evaluating Arabic diacritization models.It is built exclusively from the Tashkeela corpus for the training set and a refined version of the Fadel Tashkeela test set for the test set.
Dataset Overview
Training Data:
Source: Cleaned version of the Tashkeela corpus (original data is ~75 million words, mostly Classical… See the full description on the dataset page: https://huggingface.co/datasets/Misraj/Sadeed_Tashkeela.DIA2-Tashkeela-Gold
DIA2: A Comprehensive and Diverse Diacritized Modern Standard Arabic Corpus
DIA2 is a large-scale, natively sourced, and diacritized Modern Standard Arabic corpus
designed for NLP research and LLM development. It is curated from 28 diverse Arabic
sources including books, news articles, encyclopedic content, and poetry, and explicitly
avoids machine-translated content.
This repository contains the Tashkeela-Gold subset of DIA2.
The full DIA2 release consists of three datasets:… See the full description on the dataset page: https://huggingface.co/datasets/DIA2-Arabic/DIA2-Tashkeela-Gold.ashaar-tashkeel
Ashaar — Model-Inferred Tashkeel
Model-inferred diacritization (tashkeel) of 6,934,210 Arabic poetry verses
from the arbml/ashaar dataset.
Source
Raw verses: arbml/ashaar (241,964 poems, ~6.98M verse lines).
Diacritization model: basharalrfooh/Fine-Tashkeel
(ByT5-Large, fine-tuned on classical Arabic Tashkeela corpus).
Inference settings: FP16 on a single NVIDIA RTX 5880 Ada (48 GB),
max_new_tokens=48, max_input_len=128, greedy decoding. ~19 hours end-to-end
at ~100… See the full description on the dataset page: https://huggingface.co/datasets/mysamai/ashaar-tashkeel.tashkeel-arabic-sentencesThis dataset contains Arabic sentences extracted from the ImruQays/Alukah-Arabic dataset.
Sentences were filtered based on their 'tashkeel' (Arabic diacritics) ratio,
with a minimum ratio of 0.3 (adjustable during extraction).
Source:
The original articles were sourced from the ImruQays/Alukah-Arabic dataset on Hugging Face.
Processing:
Articles were loaded from ImruQays/Alukah-Arabic.
Each article was split into individual sentences using a regex pattern.
For each sentence, the ratio of… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/tashkeel-arabic-sentences.
