CoolFace
Datasetpublic

abuelkhair-corpus/arabic_billion_words

Abu El-Khair Corpus is an Arabic text corpus, that includes more than five million newspaper articles. It contains over a billion and a half words in total, out of which, there are about three million unique words. The corpus is encoded with two types of encoding, namely: UTF-8, and Windows CP-1256. Also it was marked with two mark-up languages, namely: SGML, and XML.

sourceHugging Faceunknownupdated 3y agoView on Hugging Face
35likes198downloads
14 commits on main
c9481463y ago

Delete legacy JSON metadata (#4)

albertvillanova
aeb210f3y ago

rename configs to config_name

lhoestq
62e7e984y ago

add dataset_info in dataset metadata

lhoestq
f4672354y ago

remove dummmy data

mariosasko
58f00c74y ago

Align more metadata with other repo types (models,spaces) (#4607)

julien-c
d3e3dd44y ago

Remove config names as yaml keys (#4367)

lhoestq
2379aa74y ago

Update datasets task tags to align tags with models (#4067)

lhoestq
eb4a5ff5y ago

Update files from the datasets library (from 1.15.0)

system
b6ed9255y ago

Update files from the datasets library (from 1.11.0)

system
34493295y ago

Update files from the datasets library (from 1.7.0)

system
a6abaaa5y ago

Update files from the datasets library (from 1.6.1)

system
793d4155y ago

Update files from the datasets library (from 1.6.0)

system
0988e705y ago

Update files from the datasets library (from 1.3.0)

system
9a071df5y ago

Update files from the datasets library (from 1.2.0)

system