CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Skylion007 /openwebtext Dataset Card for "openwebtext" Dataset Summary An open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2. This distribution was created by Aaron Gokaslan and Vanya Cohen of Brown University. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances plain_text Size of downloaded dataset… See the full description on the dataset page: https://huggingface.co/datasets/Skylion007/openwebtext.texttext-generation1M<n<10M534 likes51k downloads2mo agoHugging Face02apollo-research /Skylion007-openwebtext-tokenizer-gpt21M<n<10M3 likes2.8k downloads3y agoHugging Face03spachava /openwebtext10M<n<100M1 likes1.5k downloads3y agoHugging Face04chanind /openwebtext-gpt21M<n<10M0 likes1.4k downloads2y agoHugging Face05kenhktsui /openwebtext_quality_score_v1 Dataset Card for "openwebtext_quality_score_v1" Adding quality score v1 to Skylion007/openwebtext More Information needed texttext-generation1M<n<10M0 likes1.3k downloads3y agoHugging Face06Elriggs /openwebtext-100k Dataset Card for "openwebtext-100k" More Information needed text100K<n<1M8 likes1.2k downloads3y agoHugging Face07dylanebert /openwebtext Dataset Card for "openwebtext" Dataset Summary An open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2. This distribution was created by Aaron Gokaslan and Vanya Cohen of Brown University. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances plain_text Size of downloaded dataset files: 13.51 GB Size of the… See the full description on the dataset page: https://huggingface.co/datasets/dylanebert/openwebtext.texttext-generation1M<n<10M4 likes946 downloads1y agoHugging Face08vietgpt /the_pile_openwebtext2 Dataset Card for "the_pile_openwebtext2" More Information needed text10M<n<100M5 likes846 downloads3y agoHugging Face09chanind /openwebtext-gemma OpenWebTextCorpus tokenized for Gemma This dataset is a pre-tokenized version of the Skylion007/openwebtext dataset using the gemma tokenizer. As such, this dataset follows the same licensing as the original openwebtext dataset. This pre-tokenization is done as a performance optimization for using the openwebtext dataset with a Gemma model (gemma-2b, gemma-2b-it, gemma-7b, gemma-7b-it). This dataset was created using SAELens, with the following settings: context_size: 8192… See the full description on the dataset page: https://huggingface.co/datasets/chanind/openwebtext-gemma.1M<n<10M1 likes779 downloads2y agoHugging Face10pccl-org /Skylion007-openwebtext-tokenizer-gpt2-12810M<n<100M0 likes699 downloads2y agoHugging Face11GulkoA /openwebtext-tokenized-Llama-3.2OpenWebText dataset (open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2) tokenized for Llama 3.2 models Useful for accelerated training and testing of sparse autoencoders Context size: 128, not shuffled 10M<n<100M0 likes687 downloads2y agoHugging Face12NeelNanda /openwebtext-tokenized-9b Dataset Card for "openwebtext-tokenized-9b" More Information needed 1M<n<10M1 likes676 downloads4y agoHugging Face13chanind /openwebtext-gemma-10241M<n<10M0 likes617 downloads2y agoHugging Face14Geralt-Targaryen /openwebtext2A cleaned version of OpenWebText2 by removing non-English, duplicated, copyrighted, and low-quality (too short, too many special characters, etc) samples. This dataset has also been decontaminated with respect to the following benchmarks based on n-gram overlap: GLUE (dev set of SST-2, CoLA, QQP, WNLI, RTE, QNLI, MNLI; test set of MPRC) SIQA, PIQA, QASC, CSQA, HellaSWAG (all dev set) CONLL 2003 BLIMP MAIN BoolQ (dev set) WinoGrande (dev set) ANLI (test set) ARC easy and challenge (test set)… See the full description on the dataset page: https://huggingface.co/datasets/Geralt-Targaryen/openwebtext2.text10M<n<100M8 likes594 downloads2y agoHugging Face15heyunzhenwhat /Skylion007-openwebtext-gpt2-1024Tokenizer: gpt2dataset: '''Skylion007/openwebtext'''context_size : 1024 1M<n<10M0 likes582 downloads2y agoHugging Face16jdeschena /openwebtexttext1M<n<10M0 likes573 downloads9mo agoHugging Face17apollo-research /Skylion007-openwebtext-tokenizer-EleutherAI-gpt-neox-20b1M<n<10M0 likes521 downloads3y agoHugging Face18YangZhoumill /openwebtextpresplit Fixed OpenWebMath train/test split This is an untokenized, deterministic shuffle of open-web-math/open-web-math pinned at commit fde8ef8de2300f5e778f56261843dab89f230815. It contains the original columns without transformation. The shuffle seed is 20260904. The test set is the first 0.05% of shuffled rows (rounded to the nearest whole document); all remaining rows are training data. Exact counts and SHA-256 checksums are in split_manifest.json. texttext-generation1M<n<10M0 likes457 downloads21d agoHugging Face19vietgpt /openwebtext_en Dataset Card for "openwebtext_en" More Information needed text1M<n<10M2 likes438 downloads3y agoHugging Face20chrisjob1021 /gpt2_tokenized_concatenated_openwebtext1M<n<10M0 likes433 downloads1y agoHugging Face21pccl-org /Skylion007-openwebtext-tokenizer-gpt2-64100M<n<1B0 likes412 downloads2y agoHugging Face22RaiBP /openwebtext2-first-30-chunks-ablation-bilingualtext1M<n<10M0 likes411 downloads3y agoHugging Face23veriga /openwebtext-gemma3-tokenized-1024-activations-layer23 OpenWebText — Gemma-3-1B Hidden State Activations (Layer 23) Precomputed hidden state activations before layer 23 of Gemma-3-1B-IT for the OpenWebText dataset, tokenized with sequence length 1024. Designed for training a Titans memory layer that replaces layer 23 of Gemma 3. Dataset Structure Each example contains the inputs to layer 23: Field Shape Dtype Description activations (1024, 1152) float32 Hidden state activations (cast from bfloat16) mask(1024… See the full description on the dataset page: https://huggingface.co/datasets/veriga/openwebtext-gemma3-tokenized-1024-activations-layer23.timeseriesother10K<n<100K0 likes410 downloads4mo agoHugging Face24lsb /openwebtext-all-minilm-l6-v2-embedding Dataset Card for "openwebtext-all-minilm-l6-v2-embedding" More Information needed text1M<n<10M0 likes401 downloads4y agoHugging Face25Marlon154 /openwebtext-gemma-2-context-128 OpenWebTextCorpus tokenized for Gemma 2 with 128 context size This dataset is a pre-tokenized version of the Skylion007/openwebtext dataset using the gemma tokenizer. As such, this dataset follows the same licensing as the original openwebtext dataset. This pre-tokenization is done as a performance optimization for using the openwebtext dataset with a Gemma model (gemma-2b, gemma-2b-it, gemma-7b, gemma-7b-it). This dataset was created using SAELens, with the following settings:… See the full description on the dataset page: https://huggingface.co/datasets/Marlon154/openwebtext-gemma-2-context-128.10M<n<100M0 likes400 downloads1y agoHugging Face26dvruette /openwebtext-tokenized-16k1M<n<10M0 likes390 downloads9mo agoHugging Face27dvruette /openwebtext-tokenized-8k1M<n<10M0 likes330 downloads9mo agoHugging Face28ThomasKendrick /openwebtext Dataset Card for "openwebtext" Dataset Summary An open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2. This distribution was created by Aaron Gokaslan and Vanya Cohen of Brown University. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances plain_text Size of downloaded dataset… See the full description on the dataset page: https://huggingface.co/datasets/ThomasKendrick/openwebtext.texttext-generation1M<n<10M0 likes321 downloads24d agoHugging Face29YangZhoumill /openwebtext-presplit Fixed OpenWebText train/test split This is an untokenized, deterministic shuffle of Skylion007/openwebtext pinned at commit 79d93d786212f7344586290adb811d4ae6a1762c. It contains the original columns without transformation. The shuffle seed is 2357. The test set is the first 0.05% of shuffled rows (rounded to the nearest whole document); all remaining rows are training data. Exact counts and SHA-256 checksums are in split_manifest.json. texttext-generation1M<n<10M0 likes308 downloads21d agoHugging Face30dvruette /openwebtext-tokenized-4k1M<n<10M0 likes288 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.