CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hf-internal-testing /tokenizers-benchimage1K<n<10K0 likes19k downloads5d agoHugging Face02hf-internal-testing /tokenizers-test-data tokenizers-test-data Test and benchmark fixtures for huggingface/tokenizers, pulled on demand by the repo Makefiles (make test / make bench / make fixtures via hf download). Layout fixtures/ — multilingual + modality corpora for cross-language encode benchmarks. Organized, documented, and reproducible: see fixtures/FIXTURES.md for provenance and fixtures/fixtures_manifest.json for exact sources, pinned revisions, and sizes. Rebuild any file with… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/tokenizers-test-data.textn<1K0 likes17k downloads11d agoHugging Face03open-source-metrics /tokenizers-dependents tokenizers metrics This dataset contains metrics about the huggingface/tokenizers package. Number of repositories in the dataset: 11460 Number of packages in the dataset: 124 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 14 packages that have more than 1000 stars. There are 41… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/tokenizers-dependents.tabularn<1K0 likes2.1k downloads2y agoHugging Face04Polygl0t /tokenizers Polygl0t Tokenizers Dataset Summary This dataset contains several subsets for training multilingual tokenizers. Every subset possesses a collection of curated text samples in different languages. Supported Tasks and Leaderboards This dataset can be used for the task of text generation, specifically for training and evaluating tokenizers in multiple languages. Languages Hindi, Bengali, English, Portuguese, and Code (a mixture of 36 programming… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/tokenizers.texttext-generation1M<n<10M0 likes462 downloads7mo agoHugging Face05SlayerLab /tokenizers SlayerLab Tokenizers Normalized tokenizer artifacts collected from the contributor directories in slayerlabs/tokenizer, pinned to source commit 1a5cd2c2e4df2287b4c19b3dbf5051f5d460fdc1. The dataset contains one row per tokenizer: the 38 workshop submissions plus the canonical SlayerLab Polish 32k tokenizer by kacperwikiel. Use the Dataset Viewer to sort, filter, and compare tokenizers without navigating folders. Columns author: contributor's exact GitHub username… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/tokenizers.tabularn<1K0 likes409 downloads27d agoHugging Face06frromano /tokenizer-scratch load_data.py Dataset Summary A music dataset with audio text modality, stored in tfrecord format. Preprocessing & Augmentation Preprocessing: aggressive Augmentation: autoaugment Splits & Sampling Split strategy: leave one out Sampling: active Quality & Labeling Quality filtering: strict Labeling: pseudo label Files load_data.py — main artifact of this repository License See the… See the full description on the dataset page: https://huggingface.co/datasets/frromano/tokenizer-scratch.0 likes306 downloads27d agoHugging Face07christopher /all-tokenizerstabular100K<n<1M0 likes266 downloads9mo agoHugging Face08christopher /popular-tokenizerstabular1K<n<10K1 likes135 downloads9mo agoHugging Face09hf-internal-testing /tokenizers-bench-data0 likes134 downloads6mo agoHugging Face10Natooka /parameter-golf-sp-tokenizers Parameter Golf SP16384 — Tokenizer + Tokenized FineWeb-10B Shards SentencePiece BPE tokenizer (vocab_size=16384, byte_fallback=True) + the full FineWeb-10B corpus pre-tokenized with it. Companion artifact to the chaoscontrol submission pipeline; published to make submission-day setup frictionless — no corpus download, no re-tokenization. Files Tokenizer (root) fineweb_16384_bpe.model — SentencePiece model (455 KB). fineweb_16384_bpe.vocab — Human-readable… See the full description on the dataset page: https://huggingface.co/datasets/Natooka/parameter-golf-sp-tokenizers.10B<n<100B0 likes79 downloads5mo agoHugging Face11alexzyqi /GPT4Scene_VLN-R1_tokenizers Project Page: VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning text1M<n<10M0 likes73 downloads1y agoHugging Face12hac541309 /multilingual_tokenizersCollection of tokenizers from various sources. My own are Apache 2.0 but others are not. They are each accompanied by their license. text0 likes71 downloads3y agoHugging Face13christopher /test-tokenizerstabularn<1K0 likes41 downloads8mo agoHugging Face14open-llm-leaderboard-old /details_ewqr2130__llama_ppo_1e6_new_tokenizerstep_8000 Dataset Card for Evaluation run of ewqr2130/llama_ppo_1e6_new_tokenizerstep_8000 Dataset automatically created during the evaluation run of model ewqr2130/llama_ppo_1e6_new_tokenizerstep_8000 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ewqr2130__llama_ppo_1e6_new_tokenizerstep_8000.0 likes37 downloads3y agoHugging Face15christopher /models-for-tokenizers-metadatatabular100K<n<1M0 likes33 downloads6mo agoHugging Face16TurboPascal /tokenizers_example_zh_en用于训练分词器的基础文本 texttext-generation1M<n<10M2 likes27 downloads3y agoHugging Face17hf-internal-testing /tokenizers_test_data0 likes26 downloads11mo agoHugging Face18catherinearnett /bilingual_tokenizers20 likes21 downloads4mo agoHugging Face19catherinearnett /monolingual_tokenizers0 likes19 downloads5mo agoHugging Face20Mixture-of-tokenizers /Tokenizers-Tokens0 likes17 downloads1y agoHugging Face21cyberuser0x33 /model-tokenizers0 likes12 downloads2mo agoHugging Face22Mixture-of-tokenizers /Tokenizers-Metricstabularn<1K0 likes9 downloads1y agoHugging Face23nico-martin /tokenizers.js0 likes9 downloads1y agoHugging Face24aoUTlum /jp-datasets-for-tokenizerstext100K<n<1M0 likes8 downloads5mo agoHugging Face25catherinearnett /bilingual_tokenizers0 likes6 downloads5mo agoHugging Face26catherinearnett /trilingual-tokenizers0 likes6 downloads4mo agoHugging Face27DaveGabe /cwt-tokenizers0 likes6 downloads3mo agoHugging Face28mohit-vashishta /bhaari-tokenizers0 likes6 downloads3mo agoHugging Face29vhumenn /tokenizers0 likes5 downloads2y agoHugging Face30Sovesh /autoresearch-tokenizers0 likes5 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.