tokenizers
tokenizers-benchtokenizers-test-data
tokenizers-test-data
Test and benchmark fixtures for huggingface/tokenizers,
pulled on demand by the repo Makefiles (make test / make bench / make fixtures
via hf download).
Layout
fixtures/ — multilingual + modality corpora for cross-language encode
benchmarks. Organized, documented, and reproducible: see
fixtures/FIXTURES.md for provenance and
fixtures/fixtures_manifest.json for
exact sources, pinned revisions, and sizes. Rebuild any file with… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/tokenizers-test-data.tokenizers-dependents
tokenizers metrics
This dataset contains metrics about the huggingface/tokenizers package.
Number of repositories in the dataset: 11460
Number of packages in the dataset: 124
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 14 packages that have more than 1000 stars.
There are 41… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/tokenizers-dependents.tokenizers
Polygl0t Tokenizers
Dataset Summary
This dataset contains several subsets for training multilingual tokenizers. Every subset possesses a collection of curated text samples in different languages.
Supported Tasks and Leaderboards
This dataset can be used for the task of text generation, specifically for training and evaluating tokenizers in multiple languages.
Languages
Hindi, Bengali, English, Portuguese, and Code (a mixture of 36 programming… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/tokenizers.tokenizers
SlayerLab Tokenizers
Normalized tokenizer artifacts collected from the contributor directories in
slayerlabs/tokenizer,
pinned to source commit 1a5cd2c2e4df2287b4c19b3dbf5051f5d460fdc1.
The dataset contains one row per tokenizer: the 38 workshop submissions plus
the canonical SlayerLab Polish 32k tokenizer by kacperwikiel. Use the Dataset
Viewer to sort, filter, and compare tokenizers without navigating folders.
Columns
author: contributor's exact GitHub username… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/tokenizers.tokenizer-scratch
load_data.py
Dataset Summary
A music dataset with audio text modality, stored in tfrecord format.
Preprocessing & Augmentation
Preprocessing: aggressive
Augmentation: autoaugment
Splits & Sampling
Split strategy: leave one out
Sampling: active
Quality & Labeling
Quality filtering: strict
Labeling: pseudo label
Files
load_data.py — main artifact of this repository
License
See the… See the full description on the dataset page: https://huggingface.co/datasets/frromano/tokenizer-scratch.
