CoolFace
Datasetpublic

Grounded-Entropy/wikitext

Dataset Card for "wikitext" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License. Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over 110 times larger. The WikiText dataset also features a far… See the full description on the dataset page: https://huggingface.co/datasets/Grounded-Entropy/wikitext.

sourceHugging Facecc-by-sa-3.0updated 3mo agoView on Hugging Face
0likes14downloads
1 commits on main
03ef8d63mo ago

Duplicate from Salesforce/wikitext

Grounded-Entropy, parquet-converter