CoolFace
Datasetpublic

catherinearnett/montok

MonTok: A Suite of Monolingual Tokenizers This is a set of monolingual tokenizers for 98 languages. For each language, there are Unigram, BPE, and SuperBPE tokenizers, ranging in vocabulary size from around 6k to over 200k. Training Details Training Data All tokenizers are trained on samples of the data used to the train the Goldfish language models. The tokenizers were either trained on scaled or unscaled data. This refers to whether the models… See the full description on the dataset page: https://huggingface.co/datasets/catherinearnett/montok.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
4likes18kdownloads

catherinearnett/montok · main · files are served by the source, never re-hosted here

catherinearnett/montok · CoolFace