CoolFace
Datasetpublic

Salesforce/wikitext

Dataset Card for "wikitext" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License. Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over 110 times larger. The WikiText dataset also features a far… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/wikitext.

sourceHugging Facecc-by-sa-3.0updated 3y agoView on Hugging Face
803likes1.8mdownloads
16 commits on main
b08601e3y ago

Convert dataset to Parquet (#8)

albertvillanova
f5562963y ago

Explain difference between raw/non-raw variants (#3)

albertvillanova
dfd72873y ago

Convert dataset sizes from base 2 to base 10 in the dataset card (#6)

albertvillanova
227f3674y ago

add dataset_info in dataset metadata

lhoestq
56dfaf74y ago

remove dummmy data

mariosasko
5fd4d904y ago

Align/fix license metadata info (#4613)

julien-c
5e61f1f4y ago

Remove config names as yaml keys (#4367)

lhoestq
bc3a8024y ago

Remove a copy-paste sentence in dataset cards (#4281)

albertvillanova
dd6c0a84y ago

Update datasets task tags to align tags with models (#4067)

lhoestq
3a4a8735y ago

Update files from the datasets library (from 1.16.0)

system
e042eb05y ago

Update files from the datasets library (from 1.8.0)

system
a7ad8d85y ago

Update files from the datasets library (from 1.7.0)

system
3541efa5y ago

Update files from the datasets library (from 1.6.0)

system
267ce0b5y ago

Update files from the datasets library (from 1.4.0)

system
022482c5y ago

Update files from the datasets library (from 1.3.0)

system
f180f015y ago

Update files from the datasets library (from 1.0.0)

system