CoolFace
Datasetpublic

oserikov/arabic_billion_words

THIS IS A FORK FOR LOCAL USAGE. Abu El-Khair Corpus is an Arabic text corpus, that includes more than five million newspaper articles. It contains over a billion and a half words in total, out of which, there are about three million unique words. The corpus is encoded with two types of encoding, namely: UTF-8, and Windows CP-1256. Also it was marked with two mark-up languages, namely: SGML, and XML.

sourceHugging Faceunknownupdated 3y agoView on Hugging Face
0likes81downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
oserikov/arabic_billion_words · CoolFace