datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tokenizers-dependents
tokenizers metrics
This dataset contains metrics about the huggingface/tokenizers package.
Number of repositories in the dataset: 11460
Number of packages in the dataset: 124
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 14 packages that have more than 1000 stars.
There are 41… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/tokenizers-dependents.RULER-8192-Qwen2.5-3B-tokenizerRULER-32768-Qwen2.5-3B-tokenizerall-tokenizersRULER-32768-llama-3.1-tokenizer-chat-templatenepali-tokenizer-corpusgpt2-tokenizer-corpuspopular-tokenizersRULER-16384-llama-3.2-tokenizerRULER-4096-llama-3.2-tokenizerRULER-65536-llama-3.1-tokenizer-chat-templateRULER-131072-llama-3.1-tokenizer-chat-templateRULER-4096-llama-3.1-tokenizer-chat-templateRULER-131072-llama-3.2-tokenizerRULER-32768-llama-3.2-tokenizernepali-tokenizer-corpusRULER-8192-llama-3.1-tokenizer-chat-templatefarsi_tokenizer_robustness
TokSuite Benchmark (Farsi Collection)
Dataset Description
This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Farsi (Persian) language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness.
Curated by: R3 Research Team
Language(s): Farsi/Persian (fa)
License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/r-three/farsi_tokenizer_robustness.RULER-65536-llama-3.2-tokenizerRULER-16384-Qwen2.5-3B-tokenizerRULER-131072-Qwen2.5-3B-tokenizeroasst2_orpo_mix_tokenizer_phi_3_v1
https://huggingface.co/datasets/NickyNicky/orpo-dpo-mix-54k
RULER-4096-llama-3.1-tokenizertest-tokenizersnepali-tokenizer-corpusRULER-4096-Qwen2.5-3B-tokenizertencentdata_speech_tokenizer
Dataset Card for "tencentdata_speech_tokenizer"
More Information needed
RULER-16384-llama-3.1-tokenizer-chat-templatemodels-for-tokenizers-metadataRULER-8192-llama-3.2-tokenizer
