tashkil
Datasets
All datasets matching “tashkil”arabic_tashkil_dataset
Arabic Tashkil (Diacritization) Dataset 📖✨
Dataset Summary
This is a massive, high-quality, Gold-Standard dataset designed explicitly for training Arabic Automatic Diacritization (Tashkil) AI models (such as ByT5, AraT5, or Custom Transformers).
The dataset contains 1,494,228 heavily vocalized pages (~2.47 GB of data) extracted from Classical Arabic and Islamic texts sourced from Thahabi.org.
To ensure the highest possible ground-truth quality, every single page… See the full description on the dataset page: https://huggingface.co/datasets/freococo/arabic_tashkil_dataset.Tashkil-ASR-evalQari-OCR-0.2.2-Arabic-2B_Qari-0.2-eval-tashkilQari-OCR-0.1-VL-2B-Instruct_Qari-0.2-eval-tashkilQwen2-VL-2B-Instruct-unsloth-bnb-4bit_Qari-0.2-eval-tashkileasyocr_Qari-0.2-eval-tashkil
