datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
meta-llama-Llama-3.2-1B-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
toksuite_pretraining_dataaya-expanse-8b-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
gpt-4o-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
toksuite_english
Dataset Card for Tokenization Robustness
TokSuite Benchmark (English Collection)
Dataset Description
This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness in isolation. This specific collection contains English multiple-choice text completion questions paired with a wide range of real-world surface-form perturbations that are known to interact… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_english.toksuite_italian
Dataset Card for Tokenization Robustness
TokSuite Benchmark (Italian Collection)
Dataset Description
This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Italian language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness.
Curated by: R3 Research Team… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_italian.facebook-xglm-564M-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
toksuite_chinese
Dataset Card for Tokenization Robustness
TokSuite Benchmark (Chinese Collection)
Dataset Description
This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Chinese language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness.
Curated by: R3 Research Team… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_chinese.toksuite_turkish
Dataset Card for Tokenization Robustness
TokSuite Benchmark (Turkish Collection)
Dataset Description
This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Turkish language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness.
Curated by: R3 Research Team… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_turkish.Qwen-Qwen3-8B-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
byt5-small-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
Qwen-Qwen3-8B-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
comma-v0.1-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
toksuite_stem
Dataset Card for Tokenization Robustness
TokSuite Benchmark (STEM Collection)
Dataset Description
This dataset is the STEM subset of the TokSuite benchmark, designed to evaluate how tokenizer choice affects model behavior under realistic formatting, notation, and surface-form perturbations in technical text. TokSuite includes specialized benchmarks for mathematics and STEM, with the STEM subset containing 44 canonical technical questions paired with a… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_stem.gpt2-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
toksuite_farsi
Dataset Card for Tokenization Robustness
TokSuite Benchmark (Farsi Collection)
Dataset Description
This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Farsi (Persian) language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness.
Curated by: R3 Research… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_farsi.mistralai-tekken-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
toksuite_math
Dataset Card for Tokenization Robustness (Math)
TokSuite Benchmark (Math Collection)
Dataset Description
This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model behavior under controlled conditions.
This specific subset focuses on mathematical text completion, containing multiple-choice math questions with a variety of surface-form perturbations that stress tokenizer… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_math.toksuite_general
Dataset Card for Tokenization Robustness
TokSuite Bonus Benchmarks (General Collection)
This is a bonus TokSuite dataset containing a small set of high-signal examples that highlight surface-form variations known to affect tokenization robustness. It includes canonical questions alongside perturbations such as abbreviations, character deletion, currency symbols, diverse date formats, and unusual formatting. These examples focus on tokenization challenges that commonly… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_general.microsoft-Phi-3-mini-4k-instruct-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
tokenmonster-englishcode-32000-consistent-v1-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
bert-base-multilingual-cased-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
gemma-2b-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
bloom-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
aya-expanse-8b-toksuite-detokenizedQwen-Qwen3-8B-toksuite-detokenized
