CoolFace
Datasetpublic

toksuite/toksuite_general

Dataset Card for Tokenization Robustness TokSuite Bonus Benchmarks (General Collection) This is a bonus TokSuite dataset containing a small set of high-signal examples that highlight surface-form variations known to affect tokenization robustness. It includes canonical questions alongside perturbations such as abbreviations, character deletion, currency symbols, diverse date formats, and unusual formatting. These examples focus on tokenization challenges that… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_general.

sourceHugging Facemitupdated 8mo agoView on Hugging Face
0likes75downloads
Dataset Card

Dataset Card for Tokenization Robustness

<!-- Provide a quick summary of the dataset. -->

<img src="toksuite-logo.png" alt="TokSuite Logo" width="250px" style="margin-left:'auto' margin-right:'auto' display:'block'"/>

TokSuite Bonus Benchmarks (General Collection)

This is a bonus TokSuite dataset containing a small set of high-signal examples that highlight surface-form variations known to affect tokenization robustness. It includes canonical questions alongside perturbations such as abbreviations, character deletion, currency symbols, diverse date formats, and unusual formatting. These examples focus on tokenization challenges that commonly arise in real-world text, providing a compact complement to the main TokSuite benchmarks.

Additional Information

Dataset Curators

The dataset was curated by the TokSuite research team at R3.

Licensing Information

MIT License

Citation Information

If you use this dataset in your research, please cite the TokSuite paper:

@inproceedings{toksuite2026, title={TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior}, author={Altıntaş, Gül Sena and Ehghaghi, Malikeh and Lester, Brian and Liu, Fengyuan and Zhao, Wanru and Ciccone, Marco and Raffel, Colin}, booktitle={Preprint}, year={2026},, arxiv={https://arxiv.org/abs/2512.20757}, url={TBD} }

Paper: TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior

Contributions

This dataset is part of TokSuite, which includes:

  • —14 language models with identical architectures but different tokenizers
  • —Multilingual benchmark datasets (English, Turkish, Italian, Farsi, Chinese)
  • —Comprehensive analysis of tokenization's impact on model behavior

Contact

For questions or issues related to this dataset, please refer to the TokSuite project or contact the authors of the paper.


<div align="center">

Part of the [TokSuite Project](TBD)

Understanding Tokenization's Role in Language Model Behavior

</div>