datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikitext
Dataset Card for "wikitext"
Dataset Summary
The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified
Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License.
Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over
110 times larger. The WikiText dataset also features a far… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/wikitext.wikitext_document_level
Wikitext Document Level
This is a modified version of https://huggingface.co/datasets/wikitext that returns Wiki pages instead of Wiki text line-by-line. The original readme is contained below.
Dataset Card for "wikitext"
Dataset Summary
The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified
Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/wikitext_document_level.wikitext-2wikitext_alltime
RealTimeData Monthly Collection - Wikipedia
This datasets contains different versions of the 500 selected wikipedia articles from Wikipedia that were updated every months from 2017 to current.
To access articles in a specific month, simple run the following:
ds = datasets.load_dataset('RealTimeData/wikitext_alltime', '2020-02')
This will give you the 2020-02 version of the 500 selected wiki pages that were just updated in 2020-02.
Want to crawl the data by your own?… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/wikitext_alltime.wikitext-103-raw-v1wikitext-2-raw-v1-preprocessedoscar_wikitext_bookcorpus
Dataset Card for "oscar_wikitext_bookcorpus"
More Information needed
wikitext-103-raw-v1_gpt2-20k
Dataset Card for "wikitext-103-raw-v1_gpt2-20k"
More Information needed
wikitext__wikitext-2-raw-v1
Dataset Card for "wikitext__wikitext-2-raw-v1"
More Information needed
bergson-wikitext-512-chunks
bergson-wikitext-512-chunks
Wikitext-2 (Salesforce/wikitext, wikitext-2-raw-v1) pre-chunked into
512-GPT-2-token rows for training-data-attribution experiments with
bergson, replicating the data setup
of the MAGIC paper (Ilyas & Engstrom 2025, arXiv:2504.16430): each row is one
attribution unit / query.
train: first 4,608 chunks of the concatenated, GPT-2-tokenized
wikitext-2 train split (empty rows dropped before concatenation).
test: first 256 chunks of the wikitext-2 test… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/bergson-wikitext-512-chunks.wiki_text_embeddings
Dataset Card for "wiki_text_embeddings"
More Information needed
wiki_text
Dataset Card for "wiki_text"
More Information needed
bookcorpus-wikitext-ccnews-tinystories-chunkedbookcorpus-wikitext-ccnews-sometinystories-chunkedwikitext
Dataset Card for "wikitext"
Dataset Summary
The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified
Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License.
Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over
110 times larger. The WikiText dataset also features a far larger… See the full description on the dataset page: https://huggingface.co/datasets/LDFQ/wikitext.smol-multilingual-wikitextwikitext-tags-deberta-basewikitext_neighbors_tokenizedwikitext-103-raw-v1_sents_min_len10_max_len30_princeton-nlp_sup-simcse-roberta-largewikitext-2-raw-v1wikitext_multilingual_300kwikitext_neighbors_t5__bert-uncasewikitext-tags-modernbertdohmatob-wikitext-103wikitext_neighbors_t5__bert-uncase_w_tensorswikitext
Dataset Card for "wikitext"
Dataset Summary
The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified
Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License.
Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over
110 times larger. The WikiText dataset also features a far… See the full description on the dataset page: https://huggingface.co/datasets/HarvestEclipse/wikitext.wikitext-tags-robertaWikitext-TL39wikitext-103-raw-v1_sents_min_len10_max_len30_openai_clip-vit-base-patch32wikitext
Dataset Card for "wikitext"
Dataset Summary
The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified
Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License.
Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over
110 times larger. The WikiText dataset also features a far… See the full description on the dataset page: https://huggingface.co/datasets/TiGa-RCE/wikitext.
