datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tokenizers-benchtokenizers-test-data
tokenizers-test-data
Test and benchmark fixtures for huggingface/tokenizers,
pulled on demand by the repo Makefiles (make test / make bench / make fixtures
via hf download).
Layout
fixtures/ — multilingual + modality corpora for cross-language encode
benchmarks. Organized, documented, and reproducible: see
fixtures/FIXTURES.md for provenance and
fixtures/fixtures_manifest.json for
exact sources, pinned revisions, and sizes. Rebuild any file with… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/tokenizers-test-data.tokenizers-dependents
tokenizers metrics
This dataset contains metrics about the huggingface/tokenizers package.
Number of repositories in the dataset: 11460
Number of packages in the dataset: 124
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 14 packages that have more than 1000 stars.
There are 41… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/tokenizers-dependents.tokenizers
Polygl0t Tokenizers
Dataset Summary
This dataset contains several subsets for training multilingual tokenizers. Every subset possesses a collection of curated text samples in different languages.
Supported Tasks and Leaderboards
This dataset can be used for the task of text generation, specifically for training and evaluating tokenizers in multiple languages.
Languages
Hindi, Bengali, English, Portuguese, and Code (a mixture of 36 programming… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/tokenizers.tokenizers
SlayerLab Tokenizers
Normalized tokenizer artifacts collected from the contributor directories in
slayerlabs/tokenizer,
pinned to source commit 1a5cd2c2e4df2287b4c19b3dbf5051f5d460fdc1.
The dataset contains one row per tokenizer: the 38 workshop submissions plus
the canonical SlayerLab Polish 32k tokenizer by kacperwikiel. Use the Dataset
Viewer to sort, filter, and compare tokenizers without navigating folders.
Columns
author: contributor's exact GitHub username… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/tokenizers.tokenizer-scratch
load_data.py
Dataset Summary
A music dataset with audio text modality, stored in tfrecord format.
Preprocessing & Augmentation
Preprocessing: aggressive
Augmentation: autoaugment
Splits & Sampling
Split strategy: leave one out
Sampling: active
Quality & Labeling
Quality filtering: strict
Labeling: pseudo label
Files
load_data.py — main artifact of this repository
License
See the… See the full description on the dataset page: https://huggingface.co/datasets/frromano/tokenizer-scratch.all-tokenizerspopular-tokenizerstokenizers-bench-dataparameter-golf-sp-tokenizers
Parameter Golf SP16384 — Tokenizer + Tokenized FineWeb-10B Shards
SentencePiece BPE tokenizer (vocab_size=16384, byte_fallback=True) + the full FineWeb-10B corpus pre-tokenized with it. Companion artifact to the chaoscontrol submission pipeline; published to make submission-day setup frictionless — no corpus download, no re-tokenization.
Files
Tokenizer (root)
fineweb_16384_bpe.model — SentencePiece model (455 KB).
fineweb_16384_bpe.vocab — Human-readable… See the full description on the dataset page: https://huggingface.co/datasets/Natooka/parameter-golf-sp-tokenizers.GPT4Scene_VLN-R1_tokenizers
Project Page: VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning
multilingual_tokenizersCollection of tokenizers from various sources.
My own are Apache 2.0 but others are not.
They are each accompanied by their license.
test-tokenizersdetails_ewqr2130__llama_ppo_1e6_new_tokenizerstep_8000
Dataset Card for Evaluation run of ewqr2130/llama_ppo_1e6_new_tokenizerstep_8000
Dataset automatically created during the evaluation run of model ewqr2130/llama_ppo_1e6_new_tokenizerstep_8000 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ewqr2130__llama_ppo_1e6_new_tokenizerstep_8000.models-for-tokenizers-metadatatokenizers_example_zh_en用于训练分词器的基础文本
tokenizers_test_databilingual_tokenizers2monolingual_tokenizersTokenizers-Tokensmodel-tokenizersTokenizers-Metricstokenizers.jsjp-datasets-for-tokenizersbilingual_tokenizerstrilingual-tokenizerscwt-tokenizersbhaari-tokenizerstokenizersautoresearch-tokenizers
