CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01community-datasets /tashkeela Dataset Card for Tashkeela Dataset Summary It contains 75 million of fully vocalized words mainly 97 books from classical and modern Arabic language. Supported Tasks and Leaderboards [More Information Needed] Languages The dataset is based on Arabic. Dataset Structure Data Instances {'book':… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/tashkeela.texttext-generationn<1K6 likes221 downloads2y agoHugging Face02Misraj /Sadeed_Tashkeelagated 📚 Sadeed Tashkeela Arabic Diacritization Dataset The Sadeed dataset is a large, high-quality Arabic diacritized corpus optimized for training and evaluating Arabic diacritization models.It is built exclusively from the Tashkeela corpus for the training set and a refined version of the Fadel Tashkeela test set for the test set. Dataset Overview Training Data: Source: Cleaned version of the Tashkeela corpus (original data is ~75 million words, mostly Classical… See the full description on the dataset page: https://huggingface.co/datasets/Misraj/Sadeed_Tashkeela.texttext-generation1M<n<10M16 likes153 downloads7d agoHugging Face03DIA2-Arabic /DIA2-Tashkeela-Goldgated DIA2: A Comprehensive and Diverse Diacritized Modern Standard Arabic Corpus DIA2 is a large-scale, natively sourced, and diacritized Modern Standard Arabic corpus designed for NLP research and LLM development. It is curated from 28 diverse Arabic sources including books, news articles, encyclopedic content, and poetry, and explicitly avoids machine-translated content. This repository contains the Tashkeela-Gold subset of DIA2. The full DIA2 release consists of three datasets:… See the full description on the dataset page: https://huggingface.co/datasets/DIA2-Arabic/DIA2-Tashkeela-Gold.texttext-generation10K<n<100K0 likes27 downloads6mo agoHugging Face04mysamai /ashaar-tashkeel Ashaar — Model-Inferred Tashkeel Model-inferred diacritization (tashkeel) of 6,934,210 Arabic poetry verses from the arbml/ashaar dataset. Source Raw verses: arbml/ashaar (241,964 poems, ~6.98M verse lines). Diacritization model: basharalrfooh/Fine-Tashkeel (ByT5-Large, fine-tuned on classical Arabic Tashkeela corpus). Inference settings: FP16 on a single NVIDIA RTX 5880 Ada (48 GB), max_new_tokens=48, max_input_len=128, greedy decoding. ~19 hours end-to-end at ~100… See the full description on the dataset page: https://huggingface.co/datasets/mysamai/ashaar-tashkeel.texttext-generation1M<n<10M0 likes22 downloads5mo agoHugging Face05riotu-lab /tashkeel-arabic-sentencesThis dataset contains Arabic sentences extracted from the ImruQays/Alukah-Arabic dataset. Sentences were filtered based on their 'tashkeel' (Arabic diacritics) ratio, with a minimum ratio of 0.3 (adjustable during extraction). Source: The original articles were sourced from the ImruQays/Alukah-Arabic dataset on Hugging Face. Processing: Articles were loaded from ImruQays/Alukah-Arabic. Each article was split into individual sentences using a regex pattern. For each sentence, the ratio of… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/tashkeel-arabic-sentences.texttranslation100K<n<1M0 likes21 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.