CoolFace
Datasetpublicgated

DIA2-Arabic/DIA2-Tashkeela-Gold

DIA2: A Comprehensive and Diverse Diacritized Modern Standard Arabic Corpus DIA2 is a large-scale, natively sourced, and diacritized Modern Standard Arabic corpus designed for NLP research and LLM development. It is curated from 28 diverse Arabic sources including books, news articles, encyclopedic content, and poetry, and explicitly avoids machine-translated content. This repository contains the Tashkeela-Gold subset of DIA2. The full DIA2 release consists of three… See the full description on the dataset page: https://huggingface.co/datasets/DIA2-Arabic/DIA2-Tashkeela-Gold.

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
0likes27downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.