CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01blanchon /parler-tts_mls_eng_10k_snac_token_old Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.tabularautomatic-speech-recognition100K<n<1M1 likes991 downloads2y agoHugging Face02SakethVemula /fixed-tokenizer-morphscore-segmentstabular10M<n<100M0 likes534 downloads6mo agoHugging Face03Lyte /tokenizer-leaderboard Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): en License: mit Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed] Uses Direct Use [More… See the full description on the dataset page: https://huggingface.co/datasets/Lyte/tokenizer-leaderboard.tabularn<1K0 likes60 downloads4mo agoHugging Face04Adapting /empathetic_dialogues_with_special_tokenstabular10K<n<100K2 likes55 downloads4y agoHugging Face05LabARSS /MMLU-Pro-single-token-entropy Dataset Card for MMLU Pro with single token response entropy metadata for Mistral 24B, Phi4, Phi4-mini, Qwen2.5 3B MMLU Pro dataset with single token response entropy metadata for Mistral 24B, Phi4, Phi4-mini, Qwen2.5 3B Dataset Details Dataset Description Following up on the results from "When an LLM is apprehensive about its answers -- and when its uncertainty is justified", we measure the response entopy for MMLU Pro dataset when the model is prompted to… See the full description on the dataset page: https://huggingface.co/datasets/LabARSS/MMLU-Pro-single-token-entropy.tabular100K<n<1M0 likes42 downloads1y agoHugging Face06Circularmachines /Batch_indexing_machine_tokenstabular1M<n<10M0 likes22 downloads3y agoHugging Face07ClarusC64 /clinical-healing-trajectory-tokenization-phase-segmentation-v0.1What this dataset tests Whether a model can segment high-frequency recovery datainto interpretable healing phases. Required outputs phase_sequence phase_boundaries phase_confidence_0_100 Token labels acute_drop early_rebound consolidation_plateau oscillatory_instability secondary_drop delayed_rebound steady_ascent maladaptive_plateau recovery_lock_in Boundary format Use day indicesexampleacute_drop d0-d2 Typical failures naming phases without boundaries… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-healing-trajectory-tokenization-phase-segmentation-v0.1.tabulartext-classificationn<1K0 likes19 downloads8mo agoHugging Face08Gugu8 /Token-Efficiency token_efficiency_corpus A 2.5 GB CSV corpus teaching LLMs to minimize token usage in their outputs. Progresses from basic filler removal to expert-level nested reasoning compression. Contents verbose_output - The padded, wasteful version of the text efficient_output - The compressed, token-efficient equivalent technique - Compression strategy used subcategory - Specific variant of the technique difficulty - Tier 1 (easiest)… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Token-Efficiency.tabular1M<n<10M0 likes13 downloads2mo agoHugging Face09jr303 /lexical-fields-with-tokenstabular10K<n<100K0 likes10 downloads1y agoHugging Face10TokenBender /e5_FT_sentence_retrieval_task_Hindi_minitabular1K<n<10K0 likes7 downloads3y agoHugging Face11jedha0padavan /glassdoor-reviews-tokenizedBased on https://www.kaggle.com/datasets/davidgauthier/glassdoor-job-reviews/data Filtered by 30 top firms Addes columns with lemmatized and tokenized texts tabular100K<n<1M0 likes7 downloads1y agoHugging Face12EconomicTermDevelopments /tokenomic-drift-economics tokenomic drift Economics Dataset Dataset Description Summary Synthetic 200-row dataset for tokenomic drift measurement and computational experiments. Supported Tasks Economic analysis Cryptocurrency Economics research Computational economics Languages English (metadata and documentation) Python (code examples) Dataset Structure Data Fields id: Unique observation id epoch: Synthetic tokenomic epoch… See the full description on the dataset page: https://huggingface.co/datasets/EconomicTermDevelopments/tokenomic-drift-economics.tabulartabular-classificationn<1K0 likes7 downloads5mo agoHugging Face13hbXNov /distill_r1_qwen1p5b_math7500_soln_32k_tokenstabular1K<n<10K1 likes6 downloads2y agoHugging Face14hbXNov /r1_llama_3p1_8b_openthoughts_32k_tokenstabular10K<n<100K0 likes5 downloads2y agoHugging Face15Stepan4545 /token_risktabular10K<n<100K0 likes5 downloads5mo agoHugging Face16nsjain /single-document-tokenizedtabular100K<n<1M0 likes5 downloads5mo agoHugging Face17M-H-MARUF /bengali-tokenization-corpus Bengali Tokenization Corpus (25k Sentences) Dataset Description A balanced 25,000-sentence Bengali corpus designed for tokenization benchmarking. Dataset Summary This dataset is used in the manuscript: Quantifying the Tokenization Tax on Bengali: A Multi-Tokenizer Audit Across 25k Sentences MD. Muktadirul Haque Maruf, Israt Jahan Munny, Moutithi Sarker MouManuscript in preparation Domains Academic News Literary Colloquial Dialectal… See the full description on the dataset page: https://huggingface.co/datasets/M-H-MARUF/bengali-tokenization-corpus.tabulartext-classification10K<n<100K0 likes4 downloads4mo agoHugging Face18lavanyaashri /tokenlens-compression-benchmark TokenLens Compression Benchmark A benchmark dataset measuring LLM prompt compression quality across 100 Wikipedia articles in 5 categories, generated using TokenLens. Dataset Description This dataset contains compression quality measurements for 100 Wikipedia articles compressed at 9 different ratios (0.1 to 0.9) using extractive embedding-based compression. For each article and compression ratio, the dataset records tokens saved, semantic similarity, ROUGE-L… See the full description on the dataset page: https://huggingface.co/datasets/lavanyaashri/tokenlens-compression-benchmark.tabularn<1K0 likes3 downloads3mo agoHugging Face19newsmediabias /Bias-Tokens-CONLLgatedtabular100K<n<1M1 likes2 downloads3y agoHugging Face20loredanagaspar /hn_title_modeling_dataset_with_tokenstabular1M<n<10M0 likes1 downloads1y agoHugging Face21adamo1139 /tokenized_ds_stats_apt4tabularn<1K0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.