CoolFace
Datasetpublic

oserikov/arabic_billion_words

THIS IS A FORK FOR LOCAL USAGE. Abu El-Khair Corpus is an Arabic text corpus, that includes more than five million newspaper articles. It contains over a billion and a half words in total, out of which, there are about three million unique words. The corpus is encoded with two types of encoding, namely: UTF-8, and Windows CP-1256. Also it was marked with two mark-up languages, namely: SGML, and XML.

sourceHugging Faceunknownupdated 3y agoView on Hugging Face
0likes78downloads

oserikov/arabic_billion_words · main · files are served by the source, never re-hosted here