abuelkhair-corpus/arabic_billion_words
Abu El-Khair Corpus is an Arabic text corpus, that includes more than five million newspaper articles. It contains over a billion and a half words in total, out of which, there are about three million unique words. The corpus is encoded with two types of encoding, namely: UTF-8, and Windows CP-1256. Also it was marked with two mark-up languages, namely: SGML, and XML.
Delete legacy JSON metadata (#4)
rename configs to config_name
add dataset_info in dataset metadata
remove dummmy data
Align more metadata with other repo types (models,spaces) (#4607)
Remove config names as yaml keys (#4367)
Update datasets task tags to align tags with models (#4067)
Update files from the datasets library (from 1.15.0)
Update files from the datasets library (from 1.11.0)
Update files from the datasets library (from 1.7.0)
Update files from the datasets library (from 1.6.1)
Update files from the datasets library (from 1.6.0)
Update files from the datasets library (from 1.3.0)
Update files from the datasets library (from 1.2.0)
