CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01toksuitebackup /meta-llama-Llama-3.2-1B-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes1.2k downloads10mo agoHugging Face02toksuite /toksuite_pretraining_datatext100M<n<1B0 likes740 downloads6mo agoHugging Face03toksuitebackup /aya-expanse-8b-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes740 downloads10mo agoHugging Face04toksuitebackup /gpt-4o-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M1 likes733 downloads10mo agoHugging Face05toksuite /toksuite_english Dataset Card for Tokenization Robustness TokSuite Benchmark (English Collection) Dataset Description This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness in isolation. This specific collection contains English multiple-choice text completion questions paired with a wide range of real-world surface-form perturbations that are known to interact… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_english.textmultiple-choice1K<n<10K0 likes386 downloads8mo agoHugging Face06toksuite /toksuite_italian Dataset Card for Tokenization Robustness TokSuite Benchmark (Italian Collection) Dataset Description This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Italian language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness. Curated by: R3 Research Team… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_italian.textmultiple-choice1K<n<10K0 likes376 downloads8mo agoHugging Face07toksuitebackup /facebook-xglm-564M-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M1 likes360 downloads10mo agoHugging Face08toksuite /toksuite_chinese Dataset Card for Tokenization Robustness TokSuite Benchmark (Chinese Collection) Dataset Description This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Chinese language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness. Curated by: R3 Research Team… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_chinese.textmultiple-choicen<1K0 likes349 downloads8mo agoHugging Face09toksuite /toksuite_turkish Dataset Card for Tokenization Robustness TokSuite Benchmark (Turkish Collection) Dataset Description This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Turkish language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness. Curated by: R3 Research Team… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_turkish.textmultiple-choicen<1K0 likes342 downloads8mo agoHugging Face10toksuite /Qwen-Qwen3-8B-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes277 downloads9mo agoHugging Face11toksuitebackup /byt5-small-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular1M<n<10M0 likes262 downloads10mo agoHugging Face12toksuitebackup /Qwen-Qwen3-8B-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes240 downloads9mo agoHugging Face13toksuitebackup /comma-v0.1-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes223 downloads10mo agoHugging Face14toksuite /toksuite_stem Dataset Card for Tokenization Robustness TokSuite Benchmark (STEM Collection) Dataset Description This dataset is the STEM subset of the TokSuite benchmark, designed to evaluate how tokenizer choice affects model behavior under realistic formatting, notation, and surface-form perturbations in technical text. TokSuite includes specialized benchmarks for mathematics and STEM, with the STEM subset containing 44 canonical technical questions paired with a… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_stem.textmultiple-choicen<1K0 likes179 downloads8mo agoHugging Face15toksuitebackup /gpt2-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes177 downloads10mo agoHugging Face16toksuite /toksuite_farsi Dataset Card for Tokenization Robustness TokSuite Benchmark (Farsi Collection) Dataset Description This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Farsi (Persian) language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness. Curated by: R3 Research… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_farsi.textmultiple-choicen<1K0 likes160 downloads8mo agoHugging Face17toksuitebackup /mistralai-tekken-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes145 downloads10mo agoHugging Face18toksuite /toksuite_math Dataset Card for Tokenization Robustness (Math) TokSuite Benchmark (Math Collection) Dataset Description This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model behavior under controlled conditions. This specific subset focuses on mathematical text completion, containing multiple-choice math questions with a variety of surface-form perturbations that stress tokenizer… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_math.textmultiple-choicen<1K0 likes142 downloads8mo agoHugging Face19toksuite /toksuite_general Dataset Card for Tokenization Robustness TokSuite Bonus Benchmarks (General Collection) This is a bonus TokSuite dataset containing a small set of high-signal examples that highlight surface-form variations known to affect tokenization robustness. It includes canonical questions alongside perturbations such as abbreviations, character deletion, currency symbols, diverse date formats, and unusual formatting. These examples focus on tokenization challenges that commonly… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_general.textmultiple-choicen<1K0 likes75 downloads8mo agoHugging Face20toksuitebackup /microsoft-Phi-3-mini-4k-instruct-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M1 likes68 downloads10mo agoHugging Face21toksuitebackup /tokenmonster-englishcode-32000-consistent-v1-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes60 downloads10mo agoHugging Face22toksuitebackup /bert-base-multilingual-cased-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes56 downloads10mo agoHugging Face23toksuitebackup /gemma-2b-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes22 downloads10mo agoHugging Face24toksuitebackup /bloom-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). 0 likes4 downloads10mo agoHugging Face25gsaltintas /aya-expanse-8b-toksuite-detokenized0 likes2 downloads10mo agoHugging Face26gsaltintas /Qwen-Qwen3-8B-toksuite-detokenized0 likes2 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.