CoolFace
Datasetpublic

LTCB/enwik8

The dataset is based on the Hutter Prize (http://prize.hutter1.net) and contains the first 10^8 bytes of English Wikipedia in 2006 in XML

sourceHugging Facemitupdated 3y agoView on Hugging Face
14likes395downloads
9 commits on main
8d9ca883y ago

Delete legacy JSON metadata (#5)

albertvillanova
a3d620e3y ago

Convert missing dataset sizes from base 2 to base 10 in the dataset card (#3)

albertvillanova
6c7216b3y ago

Convert dataset sizes from base 2 to base 10 in the dataset card (#2)

albertvillanova
751abc74y ago

add dataset_info in dataset metadata

lhoestq
7fdc1234y ago

remove dummmy data

mariosasko
382b6d24y ago

Update Enwik8 broken link and information (#4950)

mtanghu
f3a6d994y ago

Align more metadata with other repo types (models,spaces) (#4607)

julien-c
e34d0404y ago

Adding dataset enwik8 (#4321)

Patrick Haller
acfa9b84y ago

initial commit

parquet-converter