CoolFace
5 results

tokenizer-robustness

Malikeh1375 /code-switching-tokenizer-robustness Code-Switching Dataset for Tokenizer Robustness Analysis Dataset Description This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models. Purpose Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.texttext-generation1K<n<10K2 likes286 downloads1y agoHugging Facer-three /farsi_tokenizer_robustness TokSuite Benchmark (Farsi Collection) Dataset Description This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Farsi (Persian) language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness. Curated by: R3 Research Team Language(s): Farsi/Persian (fa) License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/r-three/farsi_tokenizer_robustness.tabularmultiple-choicen<1K1 likes58 downloads11mo agoHugging FaceMalikeh1375 /tokenizer-robustness-mmlu Tokenizer Robustness MMLU Dataset This dataset contains MMLU-formatted questions and answers designed to test tokenizer robustness across different text formats and languages. Dataset Description The dataset consists of the same questions presented in 6 different formats, with both test (20 questions) and development (5 questions) sets: original - Standard formatted questions minor_spelling_errors - Questions with minor misspellings spoken_language - Questions in casual… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/tokenizer-robustness-mmlu.textn<1K0 likes54 downloads1y agoHugging FaceMalikeh1375 /tokenizer-robustness-math-mmlu Tokenizer Robustness Math MMLU Dataset This dataset contains MMLU-formatted questions and answers designed to test tokenizer robustness across different math notions. Dataset Description The dataset consists of the same questions presented in 6 different formats, with both test (20 questions) and development (5 questions) sets: unicode - Questions and choices in unicode format ascii - Questions and choices in ASCII format latex - Questions and choices in LaTeX format… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/tokenizer-robustness-math-mmlu.textn<1K0 likes26 downloads1y agoHugging FaceMalikeh1375 /farsi-tokenizer-robustness-mmlu Farsi Tokenizer Robustness MMLU Dataset This dataset contains MMLU-formatted questions and answers designed to test tokenizer robustness across different text formats and languages. textn<1K0 likes18 downloads1y agoHugging Face