datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
parler-tts_mls_eng_10k_snac_token_old
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.fixed-tokenizer-morphscore-segmentstokenizer-leaderboard
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): en
License: mit
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Direct Use
[More… See the full description on the dataset page: https://huggingface.co/datasets/Lyte/tokenizer-leaderboard.empathetic_dialogues_with_special_tokensMMLU-Pro-single-token-entropy
Dataset Card for MMLU Pro with single token response entropy metadata for Mistral 24B, Phi4, Phi4-mini, Qwen2.5 3B
MMLU Pro dataset with single token response entropy metadata for Mistral 24B, Phi4, Phi4-mini, Qwen2.5 3B
Dataset Details
Dataset Description
Following up on the results from "When an LLM is apprehensive about its answers -- and when its uncertainty is justified", we measure the response entopy for MMLU Pro dataset when the model is prompted to… See the full description on the dataset page: https://huggingface.co/datasets/LabARSS/MMLU-Pro-single-token-entropy.Batch_indexing_machine_tokensclinical-healing-trajectory-tokenization-phase-segmentation-v0.1What this dataset tests
Whether a model can segment high-frequency recovery datainto interpretable healing phases.
Required outputs
phase_sequence
phase_boundaries
phase_confidence_0_100
Token labels
acute_drop
early_rebound
consolidation_plateau
oscillatory_instability
secondary_drop
delayed_rebound
steady_ascent
maladaptive_plateau
recovery_lock_in
Boundary format
Use day indicesexampleacute_drop d0-d2
Typical failures
naming phases without boundaries… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-healing-trajectory-tokenization-phase-segmentation-v0.1.Token-Efficiency
token_efficiency_corpus
A 2.5 GB CSV corpus teaching LLMs to minimize token usage in their outputs.
Progresses from basic filler removal to expert-level nested reasoning compression.
Contents
verbose_output - The padded, wasteful version of the text
efficient_output - The compressed, token-efficient equivalent
technique - Compression strategy used
subcategory - Specific variant of the technique
difficulty - Tier 1 (easiest)… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Token-Efficiency.lexical-fields-with-tokense5_FT_sentence_retrieval_task_Hindi_miniglassdoor-reviews-tokenizedBased on https://www.kaggle.com/datasets/davidgauthier/glassdoor-job-reviews/data
Filtered by 30 top firms
Addes columns with lemmatized and tokenized texts
tokenomic-drift-economics
tokenomic drift Economics Dataset
Dataset Description
Summary
Synthetic 200-row dataset for tokenomic drift measurement and computational experiments.
Supported Tasks
Economic analysis
Cryptocurrency Economics research
Computational economics
Languages
English (metadata and documentation)
Python (code examples)
Dataset Structure
Data Fields
id: Unique observation id
epoch: Synthetic tokenomic epoch… See the full description on the dataset page: https://huggingface.co/datasets/EconomicTermDevelopments/tokenomic-drift-economics.distill_r1_qwen1p5b_math7500_soln_32k_tokensr1_llama_3p1_8b_openthoughts_32k_tokenstoken_risksingle-document-tokenizedbengali-tokenization-corpus
Bengali Tokenization Corpus (25k Sentences)
Dataset Description
A balanced 25,000-sentence Bengali corpus designed for tokenization benchmarking.
Dataset Summary
This dataset is used in the manuscript:
Quantifying the Tokenization Tax on Bengali: A Multi-Tokenizer Audit Across 25k Sentences
MD. Muktadirul Haque Maruf, Israt Jahan Munny, Moutithi Sarker MouManuscript in preparation
Domains
Academic
News
Literary
Colloquial
Dialectal… See the full description on the dataset page: https://huggingface.co/datasets/M-H-MARUF/bengali-tokenization-corpus.tokenlens-compression-benchmark
TokenLens Compression Benchmark
A benchmark dataset measuring LLM prompt compression quality across 100 Wikipedia articles in 5 categories, generated using TokenLens.
Dataset Description
This dataset contains compression quality measurements for 100 Wikipedia articles compressed at 9 different ratios (0.1 to 0.9) using extractive embedding-based compression. For each article and compression ratio, the dataset records tokens saved, semantic similarity, ROUGE-L… See the full description on the dataset page: https://huggingface.co/datasets/lavanyaashri/tokenlens-compression-benchmark.Bias-Tokens-CONLLhn_title_modeling_dataset_with_tokenstokenized_ds_stats_apt4
