CoolFace
Datasetpublic

r-three/farsi_tokenizer_robustness

TokSuite Benchmark (Farsi Collection) Dataset Description This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Farsi (Persian) language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness. Curated by: R3 Research Team Language(s): Farsi/Persian (fa) License:… See the full description on the dataset page: https://huggingface.co/datasets/r-three/farsi_tokenizer_robustness.

sourceHugging Facemitupdated 11mo agoView on Hugging Face
1likes56downloads
Dataset Card

<!-- Provide a quick summary of the dataset. -->

<img src="toksuite-logo.png" alt="TokSuite Logo" width="250px" style="margin-left:'auto' margin-right:'auto' display:'block'"/>

TokSuite Benchmark (Farsi Collection)

Dataset Description

This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Farsi (Persian) language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness.

  • Curated by: R3 Research Team
  • Language(s): Farsi/Persian (fa)
  • License: MIT License

Dataset Summary

TokSuite addresses a fundamental challenge in language model research: understanding how tokenization choices impact model behavior in isolation. The Farsi subset specifically measures model performance on canonical questions and various perturbations including orthographic variations, diacritics, morphological challenges, and noise commonly encountered when processing Farsi text.

Key Features:

  • 45 canonical questions covering general knowledge, geography, science, and language understanding
  • Multiple perturbation types reflecting real-world text variations in Farsi
  • Parallel structure with TokSuite benchmark (available in English, Turkish, Italian, Chinese)
  • Native speaker curation ensuring linguistic authenticity

Supported Tasks

  • Multiple-Choice Question Answering: Text completion format with 4 answer choices
  • Tokenizer Robustness Evaluation: Measuring performance degradation under various text perturbations
  • Multilingual NLP Benchmarking: Evaluating language models on Farsi text understanding

Languages

The dataset contains text in Farsi (Persian) written in Arabic script (language code: pes_Arab / fa).

Dataset Structure

Data Instances

An example from the dataset:

json
{
  "question": "رنگ آسمان",
  "choices": ["آبی است", "قرمز است", "سبز است", "زرد است"],
  "answer": 0,
  "answer_label": "A",
  "split": "test",
  "subcategories": "Canonical",
  "lang": "pes_Arab",
  "second_lang": "The color of the sky is",
  "coding_lang": "",
  "notes": "The color of the sky is",
  "id": "301",
  "set_id": 301.0,
  "variation_id": 1.0
}

Data Fields

FieldTypeDescription
questionstringThe question text in Farsi (Persian Arabic script)
choiceslist[string]Four multiple-choice answer options in Farsi
answerint64Index of the correct answer (0-3)
answer_labelstringLetter label of the correct answer (A, B, C, or D)
splitstringDataset split identifier (all entries are "test")
subcategoriesstringPerturbation category
langstringLanguage code (pes_Arab = Persian/Farsi in Arabic script)
second_langstringEnglish translation or description of the question
coding_langstringNot applicable for this dataset (empty string)
notesstringAdditional context about the question or perturbation type
idstringUnique question identifier
set_idfloat64Question set grouping identifier (ranges from 300-344)
variation_idfloat64Variation number within a question set

Dataset Creation

Curation Rationale

This dataset was created to:

  1. 1.Systematically evaluate how different tokenization strategies handle Farsi text
  2. 2.Measure robustness against real-world text perturbations specific to Farsi language
  3. 3.Support research into tokenization's impact on language model behavior
  4. 4.Provide standardized benchmarks for Farsi language models

The questions were designed to be straightforward with high baseline accuracy, allowing researchers to cleanly measure performance degradation when perturbations are applied.

Source Data

Data Collection and Processing
  • Canonical Questions: 40 baseline questions in English were created covering general knowledge topics
  • Translation: Native Farsi speakers translated questions to Persian
  • Perturbations: Each question underwent targeted perturbations designed to reflect morphological and orthographic characteristics of Farsi
  • Validation: Model-in-the-loop process ensured high baseline accuracy across 14 different tokenizers
Perturbation Categories
  1. 1.Canonical The baseline/standard form of Farsi text without any modifications, used as the reference point for comparing other perturbations.
  1. 1.Code Language Script Switching Mixing Farsi with English language (code-switching), randomly switching between Farsi and English words mid-sentence.
  1. 1.Colloquial Using informal, conversational Farsi instead of formal written language, including slang, dialectal variations, and everyday speech patterns.
  1. 1.Diacritics Presence/Absence Adding diacritical marks (vowel markings and other pronunciation indicators) that can be optionally included in Farsi text, which affects how words are read.
  1. 1.Keyboard Proximity Errors Typos caused by hitting adjacent keys on a keyboard, simulating common typing mistakes where the wrong character is typed due to finger placement.
  1. 1.Romanization Converting Farsi text to Finglish—writing Farsi words using English/Latin letters instead of Persian script.
  1. 1.Word Reordering Changing the order of words in sentences, testing whether tokenizers can handle different syntactic arrangements.
  1. 1.Word Spacing, Zero-Width Characters, Extra Space Manipulating spacing between words by adding extra spaces, removing spaces, or inserting invisible zero-width characters that affect how text is segmented.
Model Performance Comparison
model_namecanonicalarabic_keyboard_for_farsicode_language_script_switchingcolloquialdialectsequivalent_expressionskeyboard_proximity_errorsnumber_romanizationoptional_diacriticsromanizationspelled_outword_spacing_zero-width_characters_extra_space
Aya0.780.3460.7170.6610.5290.6070.4090.7440.4380.3460.4580.557
BLOOM0.7750.4480.770.60.5050.6750.5710.6690.5050.2760.5420.589
ByT50.7690.4780.7190.5910.5310.6160.5270.5680.4460.280.3370.476
Comma0.790.4710.660.6520.5230.660.5030.6170.4570.4490.2910.484
GPT-20.780.5690.6720.7390.5450.660.6160.4980.4360.2980.4490.573
GPT-4o0.750.4060.7440.6690.5040.7440.5880.7520.3750.3060.4660.544
Gemma-20.750.3750.5690.6880.4750.7120.5440.440.4310.4250.4460.5
Llama-3.20.7430.3550.6880.5870.550.6750.4990.9070.2910.3040.4290.46
Phi-30.820.480.6750.5930.5010.630.5420.5550.4930.3280.4690.593
Qwen-30.8570.4280.6430.5450.5410.590.5340.6440.4550.2520.3840.473
Tekken0.8420.4810.7430.5940.510.6970.5610.8530.4490.3180.5220.547
TokenMonster0.7140.5330.6220.6710.5210.610.5420.7280.5230.3520.3180.519
XGLM0.7570.4990.6690.5580.5220.7060.5390.6440.4620.2970.4150.559
mBERT0.7460.3770.6780.6780.5080.6590.5850.4020.5470.4140.2960.659
Who are the source data producers?

Native Farsi speakers curated and validated all questions and perturbations. The TokSuite research team at R3 designed the overall benchmark framework.

Annotations

Annotation process

Questions were manually created and translated by native speakers. Each perturbation was carefully designed to reflect authentic variations encountered in real-world Farsi text processing.

Who are the annotators?

Native Farsi speakers with expertise in linguistics and NLP, working as part of the TokSuite project.

Personal and Sensitive Information

The dataset contains only general knowledge questions and does not include any personal or sensitive information.

Considerations for Using the Data

Social Impact of Dataset

This dataset contributes to improving language technology for Farsi speakers by:

  • Enabling better understanding of tokenization challenges in Persian
  • Supporting development of more robust multilingual models
  • Providing standardized evaluation for Farsi NLP research

Discussion of Biases

  • Language variety: The dataset uses Modern Standard Persian and may not fully represent dialectal variations
  • Script focus: Only Arabic script is used; romanized versions are included as perturbations
  • Domain coverage: Questions focus on general knowledge and may not represent domain-specific language use
  • Question simplicity: Designed for high baseline accuracy, which may not reflect real-world task complexity

Other Known Limitations

  • Relatively small dataset size (designed for evaluation, not training)
  • Focus on multiple-choice format may not capture all aspects of language understanding
  • Perturbations are specific to Farsi's characteristics and findings may not generalize to all languages
  • Models evaluated were trained at ~1B parameters; results may differ at larger scales

Additional Information

Dataset Curators

The dataset was curated by the TokSuite research team at R3.

Licensing Information

MIT license

Citation Information

If you use this dataset in your research, please cite the TokSuite paper:

bibtex
@inproceedings{toksuite2026,
  title={TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior},
  author={Altıntaş, Gül Sena and Ehghaghi, Malikeh and Lester, Brian and Liu, Fengyuan and Zhao, Wanru and Ciccone, Marco and Raffel, Colin},
  booktitle={Preprint.},
  year={2026},
  url={TBD}
}

Paper: TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior

Contributions

This dataset is part of TokSuite, which includes:

  • 14 language models with identical architectures but different tokenizers
  • Multilingual benchmark datasets (English, Turkish, Italian, Farsi, Chinese)
  • Comprehensive analysis of tokenization's impact on model behavior

Contact

For questions or issues related to this dataset, please refer to the TokSuite project or contact the authors through the paper submission system.


<div align="center">

Part of the [TokSuite Project](TBD)

Understanding Tokenization's Role in Language Model Behavior

</div>