toksuite/toksuite_math
Dataset Card for Tokenization Robustness (Math) TokSuite Benchmark (Math Collection) Dataset Description This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model behavior under controlled conditions. This specific subset focuses on mathematical text completion, containing multiple-choice math questions with a variety of surface-form perturbations that stress… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_math.
Updated paths
Rename tokenizer_robustness_completion_math_canonical/test-00000-of-00001.parquet to toksuite_math_canonical/test-00000-of-00001.parquet
Rename tokenizer_robustness_completion_math_chinese/test-00000-of-00001.parquet to toksuite_math_chinese/test-00000-of-00001.parquet
Rename tokenizer_robustness_completion_math_decorative_unicode/test-00000-of-00001.parquet to toksuite_math_decorative_unicode/test-00000-of-00001.parquet
Rename tokenizer_robustness_completion_math_farsi/test-00000-of-00001.parquet to toksuite_math_farsi/test-00000-of-00001.parquet
Rename tokenizer_robustness_completion_math_italian/test-00000-of-00001.parquet to toksuite_math_italian/test-00000-of-00001.parquet
Rename tokenizer_robustness_completion_math_latex/test-00000-of-00001.parquet to toksuite_math_latex/test-00000-of-00001.parquet
Rename tokenizer_robustness_completion_math_space_removal/test-00000-of-00001.parquet to toksuite_math_space_removal/test-00000-of-00001.parquet
Rename tokenizer_robustness_completion_math_spelled_out/test-00000-of-00001.parquet to toksuite_math_spelled_out/test-00000-of-00001.parquet
Rename tokenizer_robustness_completion_math_turkish/test-00000-of-00001.parquet to toksuite_math_turkish/test-00000-of-00001.parquet
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Upload toksuite-logo.png
Uploading tokenizer_robustness_completion_math_turkish subset
Uploading tokenizer_robustness_completion_math_spelled_out subset
Uploading tokenizer_robustness_completion_math_space_removal subset
Uploading tokenizer_robustness_completion_math_latex subset
Uploading tokenizer_robustness_completion_math_italian subset
Uploading tokenizer_robustness_completion_math_farsi subset
Uploading tokenizer_robustness_completion_math_decorative_unicode subset
Uploading tokenizer_robustness_completion_math_chinese subset
Uploading tokenizer_robustness_completion_math_canonical subset
Upload README.md with huggingface_hub
Uploading tokenizer_robustness_completion_math_turkish subset
Uploading tokenizer_robustness_completion_math_spelled_out subset
Uploading tokenizer_robustness_completion_math_space_removal subset
Uploading tokenizer_robustness_completion_math_latex subset
Uploading tokenizer_robustness_completion_math_italian subset
Uploading tokenizer_robustness_completion_math_farsi subset
Uploading tokenizer_robustness_completion_math_decorative_unicode subset
Uploading tokenizer_robustness_completion_math_chinese subset
Uploading tokenizer_robustness_completion_math_canonical subset
Upload README.md with huggingface_hub
initial commit
