oserikov/arabic_billion_words_old
Dataset Card for Arabic Billion Words Corpus Dataset Summary Abu El-Khair Corpus is an Arabic text corpus, that includes more than five million newspaper articles. It contains over a billion and a half words in total, out of which, there are about three million unique words. The corpus is encoded with two types of encoding, namely: UTF-8, and Windows CP-1256. Also it was marked with two mark-up languages, namely: SGML, and XML. NB: this dataset is based on the… See the full description on the dataset page: https://huggingface.co/datasets/oserikov/arabic_billion_words_old.
This repository belongs to oserikov on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
