Polygl0t/tokenizers
Polygl0t Tokenizers Dataset Summary This dataset contains several subsets for training multilingual tokenizers. Every subset possesses a collection of curated text samples in different languages. Supported Tasks and Leaderboards This dataset can be used for the task of text generation, specifically for training and evaluating tokenizers in multiple languages. Languages Hindi, Bengali, English, Portuguese, and Code (a mixture of 36… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/tokenizers.
0462
