catherinearnett/montok
MonTok: A Suite of Monolingual Tokenizers This is a set of monolingual tokenizers for 98 languages. For each language, there are Unigram, BPE, and SuperBPE tokenizers, ranging in vocabulary size from around 6k to over 200k. Training Details Training Data All tokenizers are trained on samples of the data used to the train the Goldfish language models. The tokenizers were either trained on scaled or unscaled data. This refers to whether the models… See the full description on the dataset page: https://huggingface.co/datasets/catherinearnett/montok.
418k
