CoolFace
Datasetpublic

oserikov/arabic_billion_words_old

Dataset Card for Arabic Billion Words Corpus Dataset Summary Abu El-Khair Corpus is an Arabic text corpus, that includes more than five million newspaper articles. It contains over a billion and a half words in total, out of which, there are about three million unique words. The corpus is encoded with two types of encoding, namely: UTF-8, and Windows CP-1256. Also it was marked with two mark-up languages, namely: SGML, and XML. NB: this dataset is based on the… See the full description on the dataset page: https://huggingface.co/datasets/oserikov/arabic_billion_words_old.

sourceHugging Faceunknownupdated 3y agoView on Hugging Face
0likes33downloads
3 commits on main
4a0ee1d3y ago

add dataset to lfs

oserikov
7e146bc3y ago

initial commit

oserikov
6f065cd3y ago

initial commit

oserikov