datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tokenizers
Polygl0t Tokenizers
Dataset Summary
This dataset contains several subsets for training multilingual tokenizers. Every subset possesses a collection of curated text samples in different languages.
Supported Tasks and Leaderboards
This dataset can be used for the task of text generation, specifically for training and evaluating tokenizers in multiple languages.
Languages
Hindi, Bengali, English, Portuguese, and Code (a mixture of 36 programming… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/tokenizers.african-multilingual-tokenizer-challenge
African Multilingual Tokenizer Challenge dataset
The frozen public corpus for the African Multilingual Tokenizer Challenge. It contains one balanced multilingual train split and one balanced validation split.
Split
Per language
Total
Train
40,000
240,000
Validation
4,000
24,000
Languages are English (en), French (fr), Hausa (ha), Swahili (sw), Yoruba (yo) and Amharic (am). Official public-test and private-test text are deliberately absent from this repository.… See the full description on the dataset page: https://huggingface.co/datasets/Similoluwa/african-multilingual-tokenizer-challenge.code-switching-tokenizer-robustness
Code-Switching Dataset for Tokenizer Robustness Analysis
Dataset Description
This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models.
Purpose
Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.multilingual_tokenizer_benchmark
Multilingual Tokenizer Benchmark
More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root.
Natural language word count functions
Download spacy models
pip install ntlk spacy pygments underthesea camel-tools
python -m spacy download ko_core_news_sm
python -m spacy download ja_core_news_sm
python -m spacy download zh_core_web_sm
import nltk
nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.speakleash-tokenizer-5gb-sample
SpeakLeash tokenizer 42GB quality sample
Private tokenizer-training sample built from a stratified local SpeakLeash text snapshot on gb10.
Target: 5.0 GiB raw JSONL UTF-8 bytes after document filters and exact text dedup.
Format: data/tokenizer_sample_*.jsonl.zst, one JSON object per line with text, source, source_kind.
This is intended for tokenizer/BPE training convergence tests.
tokenizers_example_zh_en用于训练分词器的基础文本
tw-tokenizer-bench
tw-tokenizer-bench v0.1
台灣繁體中文 tokenizer 評測基準。24,294 筆 / 1,844 萬字元 ,5 個領域 subset。
定位是正式文體 ——法律、政府、百科、翻譯網頁。這是刻意的取捨,不是「台灣中文全貌」,
限制寫在下面的〈已知限制〉。
筆數
24,294
字元數
18,439,640
subset
5
每筆長度
300–1,500 字元(固定窗格)
語言
繁體中文(台灣)
Subsets
subset
筆數
字元數
平均長度
來源
law_judgment
5,000
4,386,700
877
司法院判決書
web_translated
5,000
4,178,890
836
ACE-2 翻譯繁中網頁
gov_news
5,000
3,849,201
770
政府新聞稿
encyclopedia
5,000
2,047,872
410
中文維基
law_statute
4,294
3… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-tokenizer-bench.mimi_tokenizer
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/mosh2i/mimi_tokenizer.
