CoolFace
Datasetpublicgated

DIA2-Arabic/DIA2-Tashkeela-Gold

DIA2: A Comprehensive and Diverse Diacritized Modern Standard Arabic Corpus DIA2 is a large-scale, natively sourced, and diacritized Modern Standard Arabic corpus designed for NLP research and LLM development. It is curated from 28 diverse Arabic sources including books, news articles, encyclopedic content, and poetry, and explicitly avoids machine-translated content. This repository contains the Tashkeela-Gold subset of DIA2. The full DIA2 release consists of three… See the full description on the dataset page: https://huggingface.co/datasets/DIA2-Arabic/DIA2-Tashkeela-Gold.

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
0likes28downloads
fileAliOsmTashkeela.parquet16.3 MBdownload

DIA2-Arabic/DIA2-Tashkeela-Gold · main · files are served by the source, never re-hosted here

This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.