CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Polygl0t /tokenizers Polygl0t Tokenizers Dataset Summary This dataset contains several subsets for training multilingual tokenizers. Every subset possesses a collection of curated text samples in different languages. Supported Tasks and Leaderboards This dataset can be used for the task of text generation, specifically for training and evaluating tokenizers in multiple languages. Languages Hindi, Bengali, English, Portuguese, and Code (a mixture of 36 programming… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/tokenizers.texttext-generation1M<n<10M0 likes479 downloads7mo agoHugging Face02Similoluwa /african-multilingual-tokenizer-challenge African Multilingual Tokenizer Challenge dataset The frozen public corpus for the African Multilingual Tokenizer Challenge. It contains one balanced multilingual train split and one balanced validation split. Split Per language Total Train 40,000 240,000 Validation 4,000 24,000 Languages are English (en), French (fr), Hausa (ha), Swahili (sw), Yoruba (yo) and Amharic (am). Official public-test and private-test text are deliberately absent from this repository.… See the full description on the dataset page: https://huggingface.co/datasets/Similoluwa/african-multilingual-tokenizer-challenge.texttext-generation100K<n<1M0 likes415 downloads22d agoHugging Face03Malikeh1375 /code-switching-tokenizer-robustness Code-Switching Dataset for Tokenizer Robustness Analysis Dataset Description This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models. Purpose Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.texttext-generation1K<n<10K2 likes286 downloads1y agoHugging Face04eduagarcia /multilingual_tokenizer_benchmark Multilingual Tokenizer Benchmark More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root. Natural language word count functions Download spacy models pip install ntlk spacy pygments underthesea camel-tools python -m spacy download ko_core_news_sm python -m spacy download ja_core_news_sm python -m spacy download zh_core_web_sm import nltk nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.tabulartext-generation100K<n<1M2 likes223 downloads1y agoHugging Face05kacperwikiel /speakleash-tokenizer-5gb-sample SpeakLeash tokenizer 42GB quality sample Private tokenizer-training sample built from a stratified local SpeakLeash text snapshot on gb10. Target: 5.0 GiB raw JSONL UTF-8 bytes after document filters and exact text dedup. Format: data/tokenizer_sample_*.jsonl.zst, one JSON object per line with text, source, source_kind. This is intended for tokenizer/BPE training convergence tests. texttext-generation1M<n<10M0 likes66 downloads3mo agoHugging Face06TurboPascal /tokenizers_example_zh_en用于训练分词器的基础文本 texttext-generation1M<n<10M2 likes28 downloads3y agoHugging Face07twinkle-ai /tw-tokenizer-benchgated tw-tokenizer-bench v0.1 台灣繁體中文 tokenizer 評測基準。24,294 筆 / 1,844 萬字元 ,5 個領域 subset。 定位是正式文體 ——法律、政府、百科、翻譯網頁。這是刻意的取捨,不是「台灣中文全貌」, 限制寫在下面的〈已知限制〉。 筆數 24,294 字元數 18,439,640 subset 5 每筆長度 300–1,500 字元(固定窗格) 語言 繁體中文(台灣) Subsets subset 筆數 字元數 平均長度 來源 law_judgment 5,000 4,386,700 877 司法院判決書 web_translated 5,000 4,178,890 836 ACE-2 翻譯繁中網頁 gov_news 5,000 3,849,201 770 政府新聞稿 encyclopedia 5,000 2,047,872 410 中文維基 law_statute 4,294 3… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-tokenizer-bench.texttext-generation10K<n<100K0 likes23 downloads27d agoHugging Face08mosh2i /mimi_tokenizer Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/mosh2i/mimi_tokenizer.texttext-generationn<1K0 likes20 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.