Salesforce/wikitext
Dataset Card for "wikitext" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License. Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over 110 times larger. The WikiText dataset also features a far… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/wikitext.
Convert dataset to Parquet (#8)
Explain difference between raw/non-raw variants (#3)
Convert dataset sizes from base 2 to base 10 in the dataset card (#6)
add dataset_info in dataset metadata
remove dummmy data
Align/fix license metadata info (#4613)
Remove config names as yaml keys (#4367)
Remove a copy-paste sentence in dataset cards (#4281)
Update datasets task tags to align tags with models (#4067)
Update files from the datasets library (from 1.16.0)
Update files from the datasets library (from 1.8.0)
Update files from the datasets library (from 1.7.0)
Update files from the datasets library (from 1.6.0)
Update files from the datasets library (from 1.4.0)
Update files from the datasets library (from 1.3.0)
Update files from the datasets library (from 1.0.0)
