toksuite/toksuite_general
Dataset Card for Tokenization Robustness TokSuite Bonus Benchmarks (General Collection) This is a bonus TokSuite dataset containing a small set of high-signal examples that highlight surface-form variations known to affect tokenization robustness. It includes canonical questions alongside perturbations such as abbreviations, character deletion, currency symbols, diverse date formats, and unusual formatting. These examples focus on tokenization challenges that… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_general.
Dataset Card for Tokenization Robustness
<!-- Provide a quick summary of the dataset. -->
<img src="toksuite-logo.png" alt="TokSuite Logo" width="250px" style="margin-left:'auto' margin-right:'auto' display:'block'"/>
TokSuite Bonus Benchmarks (General Collection)
This is a bonus TokSuite dataset containing a small set of high-signal examples that highlight surface-form variations known to affect tokenization robustness. It includes canonical questions alongside perturbations such as abbreviations, character deletion, currency symbols, diverse date formats, and unusual formatting. These examples focus on tokenization challenges that commonly arise in real-world text, providing a compact complement to the main TokSuite benchmarks.
Additional Information
Dataset Curators
The dataset was curated by the TokSuite research team at R3.
Licensing Information
MIT License
Citation Information
If you use this dataset in your research, please cite the TokSuite paper:
@inproceedings{toksuite2026, title={TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior}, author={Altıntaş, Gül Sena and Ehghaghi, Malikeh and Lester, Brian and Liu, Fengyuan and Zhao, Wanru and Ciccone, Marco and Raffel, Colin}, booktitle={Preprint}, year={2026},, arxiv={https://arxiv.org/abs/2512.20757}, url={TBD} }
Paper: TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior
Contributions
This dataset is part of TokSuite, which includes:
- 14 language models with identical architectures but different tokenizers
- Multilingual benchmark datasets (English, Turkish, Italian, Farsi, Chinese)
- Comprehensive analysis of tokenization's impact on model behavior
Contact
For questions or issues related to this dataset, please refer to the TokSuite project or contact the authors of the paper.
<div align="center">
Part of the [TokSuite Project](TBD)
Understanding Tokenization's Role in Language Model Behavior
</div>
