CoolFace
Datasetpublic

toksuite/toksuite_farsi

Dataset Card for Tokenization Robustness TokSuite Benchmark (Farsi Collection) Dataset Description This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Farsi (Persian) language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness. Curated by: R3… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_farsi.

sourceHugging Facemitupdated 8mo agoView on Hugging Face
0likes169downloads
Dataset Card

Dataset Card for Tokenization Robustness

<!-- Provide a quick summary of the dataset. -->

<img src="toksuite-logo.png" alt="TokSuite Logo" width="250px" style="margin-left:'auto' margin-right:'auto' display:'block'"/>

TokSuite Benchmark (Farsi Collection)

Dataset Description

This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Farsi (Persian) language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness.

  • Curated by: R3 Research Team
  • Language(s): Farsi/Persian (fa)
  • License: MIT License

Dataset Summary

TokSuite addresses a fundamental challenge in language model research: understanding how tokenization choices impact model behavior in isolation. The Farsi subset specifically measures model performance on canonical questions and various perturbations including orthographic variations, diacritics, morphological challenges, and noise commonly encountered when processing Farsi text.

Key Features:

  • 45 canonical questions covering general knowledge, geography, science, and language understanding
  • Multiple perturbation types reflecting real-world text variations in Farsi
  • Parallel structure with TokSuite benchmark (available in English, Turkish, Italian, Chinese)
  • Native speaker curation ensuring linguistic authenticity

Supported Tasks

  • Multiple-Choice Question Answering: Text completion format with 4 answer choices
  • Tokenizer Robustness Evaluation: Measuring performance degradation under various text perturbations
  • Multilingual NLP Benchmarking: Evaluating language models on Farsi text understanding

Languages

The dataset contains text in Farsi (Persian) written in Arabic script (language code: pes_Arab / fa).

Dataset Structure

Data Fields

FieldTypeDescription
questionstringThe question text in Farsi (Persian Arabic script)
choiceslist[string]Four multiple-choice answer options in Farsi
answerint64Index of the correct answer (0-3)
answer_labelstringLetter label of the correct answer (A, B, C, or D)
splitstringDataset split identifier (all entries are "test")
subcategoriesstringPerturbation category
langstringLanguage code (pes_Arab = Persian/Farsi in Arabic script)
second_langstringEnglish translation or description of the question
notesstringAdditional context about the question or perturbation type
idstringUnique question identifier
set_idfloat64Question set grouping identifier (ranges from 300-344)
variation_idfloat64Variation number within a question set
vanilla_cos_sim_to_canonicaldict[string, float]Cosine similarity scores between the tokenized representation of this example and its canonical form, computed using vanilla (untrimmed) token sequences for each tokenizer or model listed as keys.
trimmed_cos_sim_to_canonicaldict[string, float]Cosine similarity scores between this example and its canonical form after trimming or normalizing token sequences (e.g., removing special tokens), reported per tokenizer or model.
token_countsdict[string, integer]The number of tokens produced by each tokenizer or model when encoding the question text, used to analyze tokenization efficiency and fragmentation across tokenizers.

Dataset Creation

Curation Rationale

This dataset was created to:

  1. 1.Systematically evaluate how different tokenization strategies handle Farsi text
  2. 2.Measure robustness against real-world text perturbations specific to the Farsi language
  3. 3.Support research into the impact of tokenization on language model behavior
  4. 4.Provide standardized benchmarks for Farsi language models

The questions were designed to be straightforward with high baseline accuracy, allowing researchers to cleanly measure performance degradation when perturbations are applied.

Source Data

Data Collection and Processing
  • Canonical Questions: 40 baseline questions in English were created covering general knowledge topics
  • Translation: Native Farsi speakers translated questions to Persian
  • Perturbations: Each question underwent targeted perturbations designed to reflect morphological and orthographic characteristics of Farsi
  • Validation: Model-in-the-loop process ensured high baseline accuracy across 14 different tokenizers
Perturbation Categories
  1. 1.Canonical The baseline/standard form of Farsi text without any modifications, used as the reference point for comparing other perturbations.
  1. 1.Code/Language/Script Switching Mixing Farsi with English language (code-switching), randomly switching between Farsi and English words mid-sentence.
  1. 1.Colloquial Using informal, conversational Farsi instead of formal written language, including slang and everyday speech patterns.
  1. 1.Optional Diacritics Adding diacritical marks (vowel markings and other pronunciation indicators) that can be optionally included in Farsi text, which affects how words are read.
  1. 1.Keyboard Proximity Errors Typos caused by hitting adjacent keys on a keyboard, simulating common typing mistakes where the wrong character is typed due to finger placement.
  1. 1.Romanization Converting Farsi text to Finglish—writing Farsi words using English/Latin letters instead of Persian script.
  1. 1.Spelled-Out Forms Replaces symbols, abbreviations, or compact forms with fully spelled-out equivalents (e.g., numbers written in words). This tests tokenizer sensitivity to length expansion and lexical restructuring.
  1. 1.Word Spacing, Zero-Width Characters, Extra Space Manipulating spacing between words by adding extra spaces, removing spaces, or inserting invisible zero-width characters that affect how text is segmented.
  1. 1.Arabic Keyboard for Farsi Simulating text produced when users type Persian using an Arabic keyboard layout. This introduces systematic character substitutions (e.g., different forms of Yeh or Kaf) that preserve semantics but alter Unicode representations, stressing tokenizer sensitivity to script-level variations.
  1. 1.Dialectal Variations Introduces regional Persian dialect (e.g., Isfahani, Araki, Sorkheyi, Shirazi, Dezfouli, Kashani, Sabzevari, Mazandarani, and Kermani) forms that differ lexically or morphologically from standard Persian. These variations preserve meaning but alter surface forms, testing tokenizer generalization across dialects.
  1. 1.Equivalent Expression Replaces canonical expressions with alternative phrasings that convey the same meaning using different words or constructions. This perturbation isolates tokenizer sensitivity to paraphrasing without changing semantics.
  1. 1.Number Romanization Replaces Persian or Arabic numerals with Romanized (Latin-script) number forms (e.g., “3” → “۳”). This tests how tokenizers handle cross-script numeric representations.
Who are the source data producers?

Native Farsi speakers curated and validated all questions and perturbations. The TokSuite research team at R3 designed the overall benchmark framework.

Annotations

Annotation process

Questions were manually created and translated by native speakers. Each perturbation was carefully designed to reflect authentic variations encountered in real-world Farsi text processing.

Who are the annotators?

Native Farsi speakers with expertise in linguistics and NLP, working as part of the TokSuite project.

Personal and Sensitive Information

The dataset contains only general knowledge questions and does not include any personal or sensitive information.

Considerations for Using the Data

Social Impact of Dataset

This dataset contributes to improving language technology for Farsi speakers by:

  • Enabling better understanding of tokenization challenges in Persian
  • Supporting development of more robust multilingual models
  • Providing standardized evaluation for Farsi NLP research

Discussion of Biases

  • Language variety: The dataset uses Modern Standard Persian and may not fully represent dialectal variations
  • Script focus: Only Arabic script is used; romanized versions are included as perturbations
  • Domain coverage: Questions focus on general knowledge and may not represent domain-specific language use
  • Question simplicity: Designed for high baseline accuracy, which may not reflect real-world task complexity

Other Known Limitations

  • Relatively small dataset size (designed for evaluation, not training)
  • Focus on multiple-choice format may not capture all aspects of language understanding
  • Perturbations are specific to Farsi's characteristics and findings may not generalize to all languages
  • Models evaluated were trained at ~1B parameters; results may differ at larger scales

Additional Information

Dataset Curators

The dataset was curated by the TokSuite research team at R3.

Licensing Information

MIT license

Citation Information

If you use this dataset in your research, please cite the TokSuite paper:

bibtex
@inproceedings{toksuite2026,
  title={TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior},
  author={Altıntaş, Gül Sena and Ehghaghi, Malikeh and Lester, Brian and Liu, Fengyuan and Zhao, Wanru and Ciccone, Marco and Raffel, Colin},
  booktitle={Preprint.},
  year={2026},,
  arxiv={https://arxiv.org/abs/2512.20757},
  url={TBD}
}

Paper: TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior

Contributions

This dataset is part of TokSuite, which includes:

  • 14 language models with identical architectures but different tokenizers
  • Multilingual benchmark datasets (English, Turkish, Italian, Farsi, Chinese)
  • Comprehensive analysis of tokenization's impact on model behavior

Contact

For questions or issues related to this dataset, please refer to the TokSuite project or contact the authors of the paper.


<div align="center">

Part of the [TokSuite Project](TBD)

Understanding Tokenization's Role in Language Model Behavior

</div>