CoolFace
Datasetpublic

toksuite/toksuite_farsi

Dataset Card for Tokenization Robustness TokSuite Benchmark (Farsi Collection) Dataset Description This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Farsi (Persian) language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness. Curated by: R3… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_farsi.

sourceHugging Facemitupdated 8mo agoView on Hugging Face
0likes134downloads
settings

This repository belongs to toksuite on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nametoksuite_farsi
visibilitypublic
licencemit
gatedno
ownertoksuite
Account settings