CoolFace
Datasetpublic

oserikov/arabic_billion_words

THIS IS A FORK FOR LOCAL USAGE. Abu El-Khair Corpus is an Arabic text corpus, that includes more than five million newspaper articles. It contains over a billion and a half words in total, out of which, there are about three million unique words. The corpus is encoded with two types of encoding, namely: UTF-8, and Windows CP-1256. Also it was marked with two mark-up languages, namely: SGML, and XML.

sourceHugging Faceunknownupdated 3y agoView on Hugging Face
0likes78downloads
3 commits on main
7c83a9c3y ago

Update arabic_billion_words.py

oserikov
058ccc73y ago

initial commit

oserikov
bd53e233y ago

initial commit

oserikov