CoolFace
Datasetpublic

MohamedRashad/arabic-billion-words

Arabic Billion Words Dataset 🌕 The Abu El-Khair Arabic News Corpus (arabic-billion-words) is a comprehensive collection of Arabic text, encompassing over five million newspaper articles. The corpus is rich in linguistic diversity, containing more than a billion and a half words, with approximately three million unique words. The text is encoded in two formats: UTF-8 and Windows CP-1256, and marked up using two markup languages: SGML and XML. Data Example An… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-billion-words.

sourceHugging Faceupdated 3y agoView on Hugging Face
12likes244downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
MohamedRashad/arabic-billion-words · CoolFace