tokenizer-robustness
code-switching-tokenizer-robustness
Code-Switching Dataset for Tokenizer Robustness Analysis
Dataset Description
This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models.
Purpose
Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.farsi_tokenizer_robustness
TokSuite Benchmark (Farsi Collection)
Dataset Description
This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Farsi (Persian) language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness.
Curated by: R3 Research Team
Language(s): Farsi/Persian (fa)
License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/r-three/farsi_tokenizer_robustness.tokenizer-robustness-mmlu
Tokenizer Robustness MMLU Dataset
This dataset contains MMLU-formatted questions and answers designed to test tokenizer robustness across different text formats and languages.
Dataset Description
The dataset consists of the same questions presented in 6 different formats, with both test (20 questions) and development (5 questions) sets:
original - Standard formatted questions
minor_spelling_errors - Questions with minor misspellings
spoken_language - Questions in casual… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/tokenizer-robustness-mmlu.tokenizer-robustness-math-mmlu
Tokenizer Robustness Math MMLU Dataset
This dataset contains MMLU-formatted questions and answers designed to test tokenizer robustness across different math notions.
Dataset Description
The dataset consists of the same questions presented in 6 different formats, with both test (20 questions) and development (5 questions) sets:
unicode - Questions and choices in unicode format
ascii - Questions and choices in ASCII format
latex - Questions and choices in LaTeX format… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/tokenizer-robustness-math-mmlu.farsi-tokenizer-robustness-mmlu
Farsi Tokenizer Robustness MMLU Dataset
This dataset contains MMLU-formatted questions and answers designed to test tokenizer robustness across different text formats and languages.
