CoolFace
Datasetpublic

Abzalbek89/kk-tokenizer-fertility-baseline

Kazakh Tokenizer Fertility Baseline Reproducible fertility benchmark of subword tokenizers on the Kazakh language. Companion artifact for the paper "Tokenizer Optimization for Kazakh Small Language Models" (in preparation, target: ACM TALLIP). Headline numbers Tokenizer Fertility 🥇 Best overall kk-bpe-32k 1.679 🚨 Worst GPT-4 (cl100k) 5.895 GPT-4 penalty GPT-4 (cl100k) is 3.51× worse than the best Kazakh-trained tokenizer → The custom Kazakh… See the full description on the dataset page: https://huggingface.co/datasets/Abzalbek89/kk-tokenizer-fertility-baseline.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes601downloads

Abzalbek89/kk-tokenizer-fertility-baseline · main · files are served by the source, never re-hosted here