CoolFace
Datasetpublic

MohamedRashad/arabic-billion-words

Arabic Billion Words Dataset 🌕 The Abu El-Khair Arabic News Corpus (arabic-billion-words) is a comprehensive collection of Arabic text, encompassing over five million newspaper articles. The corpus is rich in linguistic diversity, containing more than a billion and a half words, with approximately three million unique words. The text is encoded in two formats: UTF-8 and Windows CP-1256, and marked up using two markup languages: SGML and XML. Data Example An… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-billion-words.

sourceHugging Faceupdated 3y agoView on Hugging Face
12likes244downloads

MohamedRashad/arabic-billion-words · main · files are served by the source, never re-hosted here