CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Salesforce /wikitext Dataset Card for "wikitext" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License. Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over 110 times larger. The WikiText dataset also features a far… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/wikitext.texttext-generation1M<n<10M812 likes1.9m downloads3y agoHugging Face02EleutherAI /wikitext_document_level Wikitext Document Level This is a modified version of https://huggingface.co/datasets/wikitext that returns Wiki pages instead of Wiki text line-by-line. The original readme is contained below. Dataset Card for "wikitext" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/wikitext_document_level.text10K<n<100K18 likes80k downloads2y agoHugging Face03mikasenghaas /wikitext-2text10K<n<100K4 likes3.3k downloads2y agoHugging Face04RealTimeData /wikitext_alltime RealTimeData Monthly Collection - Wikipedia This datasets contains different versions of the 500 selected wikipedia articles from Wikipedia that were updated every months from 2017 to current. To access articles in a specific month, simple run the following: ds = datasets.load_dataset('RealTimeData/wikitext_alltime', '2020-02') This will give you the 2020-02 version of the 500 selected wiki pages that were just updated in 2020-02. Want to crawl the data by your own?… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/wikitext_alltime.text10K<n<100K3 likes1.2k downloads1y agoHugging Face05iohadrubin /wikitext-103-raw-v1text10K<n<100K10 likes817 downloads4y agoHugging Face06Self-GRIT /wikitext-2-raw-v1-preprocessedtext10K<n<100K2 likes778 downloads2y agoHugging Face07AshtonIsNotHere /oscar_wikitext_bookcorpus Dataset Card for "oscar_wikitext_bookcorpus" More Information needed text10M<n<100M0 likes561 downloads3y agoHugging Face08pietrolesci /wikitext-103-raw-v1_gpt2-20k Dataset Card for "wikitext-103-raw-v1_gpt2-20k" More Information needed tabular1M<n<10M0 likes526 downloads3y agoHugging Face09closji /wikitext__wikitext-2-raw-v1 Dataset Card for "wikitext__wikitext-2-raw-v1" More Information needed text10K<n<100K1 likes407 downloads4y agoHugging Face10EleutherAI /bergson-wikitext-512-chunks bergson-wikitext-512-chunks Wikitext-2 (Salesforce/wikitext, wikitext-2-raw-v1) pre-chunked into 512-GPT-2-token rows for training-data-attribution experiments with bergson, replicating the data setup of the MAGIC paper (Ilyas & Engstrom 2025, arXiv:2504.16430): each row is one attribution unit / query. train: first 4,608 chunks of the concatenated, GPT-2-tokenized wikitext-2 train split (empty rows dropped before concatenation). test: first 256 chunks of the wikitext-2 test… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/bergson-wikitext-512-chunks.texttext-generation1K<n<10K0 likes374 downloads3mo agoHugging Face11yoandrey /wiki_text_embeddings Dataset Card for "wiki_text_embeddings" More Information needed text10M<n<100M0 likes264 downloads3y agoHugging Face12yoandrey /wiki_text Dataset Card for "wiki_text" More Information needed text10M<n<100M0 likes262 downloads3y agoHugging Face13ibm-aimc /bookcorpus-wikitext-ccnews-tinystories-chunkedtext1M<n<10M1 likes234 downloads2y agoHugging Face14ibm-aimc /bookcorpus-wikitext-ccnews-sometinystories-chunkedtext1M<n<10M0 likes228 downloads2y agoHugging Face15LDFQ /wikitext Dataset Card for "wikitext" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License. Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over 110 times larger. The WikiText dataset also features a far larger… See the full description on the dataset page: https://huggingface.co/datasets/LDFQ/wikitext.texttext-generation1M<n<10M0 likes196 downloads5mo agoHugging Face16haryoaw /smol-multilingual-wikitexttext100K<n<1M1 likes180 downloads1y agoHugging Face17hriaz /wikitext-tags-deberta-base1M<n<10M0 likes177 downloads1y agoHugging Face18iohadrubin /wikitext_neighbors_tokenized10K<n<100K0 likes165 downloads4y agoHugging Face19closji /wikitext-103-raw-v1_sents_min_len10_max_len30_princeton-nlp_sup-simcse-roberta-largetext1M<n<10M0 likes156 downloads4y agoHugging Face20KrisMinchev /wikitext-2-raw-v1text10K<n<100K0 likes145 downloads11mo agoHugging Face21haryoaw /wikitext_multilingual_300ktext1M<n<10M0 likes140 downloads1y agoHugging Face22iohadrubin /wikitext_neighbors_t5__bert-uncasetext1M<n<10M0 likes138 downloads4y agoHugging Face23hriaz /wikitext-tags-modernberttext1M<n<10M0 likes138 downloads1y agoHugging Face24KrisMinchev /dohmatob-wikitext-103text1M<n<10M0 likes125 downloads9mo agoHugging Face25iohadrubin /wikitext_neighbors_t5__bert-uncase_w_tensorstext1M<n<10M0 likes121 downloads4y agoHugging Face26HarvestEclipse /wikitext Dataset Card for "wikitext" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License. Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over 110 times larger. The WikiText dataset also features a far… See the full description on the dataset page: https://huggingface.co/datasets/HarvestEclipse/wikitext.texttext-generation1M<n<10M1 likes120 downloads11d agoHugging Face27hriaz /wikitext-tags-roberta1M<n<10M0 likes118 downloads1y agoHugging Face28linkanjarad /Wikitext-TL39texttext-generation1M<n<10M2 likes111 downloads3y agoHugging Face29closji /wikitext-103-raw-v1_sents_min_len10_max_len30_openai_clip-vit-base-patch32text1M<n<10M0 likes110 downloads4y agoHugging Face30TiGa-RCE /wikitext Dataset Card for "wikitext" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License. Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over 110 times larger. The WikiText dataset also features a far… See the full description on the dataset page: https://huggingface.co/datasets/TiGa-RCE/wikitext.texttext-generation1M<n<10M0 likes109 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.