datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
token-counts
Marin Token Counts
Token counts for all datasets used in Marin pretraining runs.
Schema
Column
Type
Description
dataset
string
Dataset identifier
marin_tokens
int
Number of tokens after tokenization
category
string
Content domain (web, code, math, academic, books, etc.)
synthetic
bool
Whether the data is LLM-generated or LLM-translated
Categories
web — Quality-classified Common Crawl text (Nemotron-CC)
code — Source code and… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/token-counts.parler-tts_mls_eng_10k_snac_token_old
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.fixed-tokenizer-morphscore-segmentsablation_tokensprompts_under_512_tokens
Under 512 Tokens Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use
Dataset Overview
Specialized collection of short-form English prompts (under 512 tokens), perfect for training models with context length constraints or faster iteration cycles.
📊 Dataset Statistics
Metric
Value
Total Files
200
Rows Per File
10,000
Total Rows
2,000,000
Token Range
1 to 511 tokens… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/prompts_under_512_tokens.globalise_NER_token_classification_dataset
Dataset Card for Dataset Name
The globalise_NER_token_classification dataset is a fine-grained dataset for the training of token-classification NER models on Dutch East-India Company texts (17th to 18th century).
Dataset Details
Dataset Description
The dataset provides 15 fine-grained labels detailing activities and people of the Dutch East-India Company (VOC), and can be used to train NER token-classification models for the
period 17th-18th century and the… See the full description on the dataset page: https://huggingface.co/datasets/globalise/globalise_NER_token_classification_dataset.tokenizer-leaderboard
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): en
License: mit
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Direct Use
[More… See the full description on the dataset page: https://huggingface.co/datasets/Lyte/tokenizer-leaderboard.empathetic_dialogues_with_special_tokensclassification_token_propagandawebnlg_tokensmedium_512_1k_tokens_prompts
Medium 512-1K Tokens Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ By using this dataset you agree to our Terms of Use.
Overview
703 high-quality English prompts whose length lies between 512 and 1 000 tokens.Every prompt has been de-duplicated, cleaned and token-counted with the GPT-2 tokenizer.
Statistics
Rows
Token range
File size
Format
703
512 – 1 000
2.9 MB
CSV
Use-cases
Medium-context language-model fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/medium_512_1k_tokens_prompts.MMLU-Pro-single-token-entropy
Dataset Card for MMLU Pro with single token response entropy metadata for Mistral 24B, Phi4, Phi4-mini, Qwen2.5 3B
MMLU Pro dataset with single token response entropy metadata for Mistral 24B, Phi4, Phi4-mini, Qwen2.5 3B
Dataset Details
Dataset Description
Following up on the results from "When an LLM is apprehensive about its answers -- and when its uncertainty is justified", we measure the response entopy for MMLU Pro dataset when the model is prompted to… See the full description on the dataset page: https://huggingface.co/datasets/LabARSS/MMLU-Pro-single-token-entropy.long_over_1k_tokens_prompts
Long Over 1K Tokens Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use
Dataset Overview
Specialized collection of long-form English prompts (≥ 1 000 tokens) for training advanced models that require extensive context and complex reasoning.
📊 Dataset Statistics
Metric
Value
Total Rows
289
Token Range
1 001 – 10 000 tokens
File Size
≈ 3.7 MB
Format
Single CSV file
Target… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/long_over_1k_tokens_prompts.PubChem10M_SELFIES_TokenizedCustom cl100k tokenized version of PubChem10M_SELFIES.
SimpleDC
simpledc-dataset
Official huggingface dataset for the SimpleDC (Simple Digestive Cancer) dataset
Please cite as:
@article{rahman2024health,
title={Health Text Simplification: An Annotated Corpus for Digestive Cancer Education and Novel Strategies for Reinforcement Learning},
author={Rahman, Md Mushfiqur and Irbaz, Mohammad Sabik and North, Kai and Williams, Michelle S and Zampieri, Marcos and Lybarger, Kevin},
journal={arXiv preprint arXiv:2401.15043},
year={2024}
}
tokenized-IELTS-writing-task-2-evaluation-DialoGPT-mediumToken_Optimization_Org
AI Safety & Bias Evaluation Conversations
Dataset Summary
This dataset contains simulated multi-turn conversations designed to evaluate AI language model behavior across two safety-critical domains: self-harm response handling and political bias. Each row represents a single evaluation scenario where an AI model's responses are assessed for safety compliance or neutrality. The dataset is intended to support research and development of safer, less biased AI systems.
All… See the full description on the dataset page: https://huggingface.co/datasets/token-opt-org/Token_Optimization_Org.token-budgets-catalog
Token Budgets — Empirical catalogue and inter-rater reliability data
Data for:
Token Budgets: An Empirical Catalog of 63 LLM-Agent Budget-Overrun
Incidents, with an Affine-Typed Rust Mitigation as a Case Study.
Sajjad Khan, 2026.
arXiv:2606.04056 — preprint.
This dataset bundles three artefacts referenced in the paper:
catalogue — the harvested catalogue of LLM-agent budget-overrun
incidents across 21 orchestration frameworks (2023–2026), 167 rows
total. The IRR-included… See the full description on the dataset page: https://huggingface.co/datasets/sajjadanwar0/token-budgets-catalog.Dynamic-Topic-RedPajama-Data-1T-100k-SubSample-max-1k-tokens
Dynamic Topic Modeling Dataset: RedPajama-1T SubSample (100k samples, 1k tokens)
📝Check out the Blog Post
This dataset represents a curated subset of the RedPajama-1T Sample dataset, specifically processed for dynamic topic modeling applications. It contains 100,000
samples from the original dataset, with each document limited to the first 1,024 tokens for consistent processing.
Dataset Overview
Name:… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/Dynamic-Topic-RedPajama-Data-1T-100k-SubSample-max-1k-tokens.whisper_transcriptions_token_idsBatch_indexing_machine_tokensrams-no-special-tokensclinical-healing-trajectory-tokenization-phase-segmentation-v0.1What this dataset tests
Whether a model can segment high-frequency recovery datainto interpretable healing phases.
Required outputs
phase_sequence
phase_boundaries
phase_confidence_0_100
Token labels
acute_drop
early_rebound
consolidation_plateau
oscillatory_instability
secondary_drop
delayed_rebound
steady_ascent
maladaptive_plateau
recovery_lock_in
Boundary format
Use day indicesexampleacute_drop d0-d2
Typical failures
naming phases without boundaries… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-healing-trajectory-tokenization-phase-segmentation-v0.1.sentence_retrieval_hindi_SFTe5_FT_sentence_retrieval_task_Hinditokenized_datasetQAtoken-optimization
AI Safety & Bias Evaluation Conversations
Dataset Summary
This dataset contains simulated multi-turn conversations designed to evaluate AI language model behavior across two safety-critical domains: self-harm response handling and political bias. Each row represents a single evaluation scenario where an AI model's responses are assessed for safety compliance or neutrality. The dataset is intended to support research and development of safer, less biased AI… See the full description on the dataset page: https://huggingface.co/datasets/CentificAIResearch/token-optimization.GTE-ModernBERT-RedPajama-Data-1T-100k-SubSample-max-1k-tokensToken-Efficiency
token_efficiency_corpus
A 2.5 GB CSV corpus teaching LLMs to minimize token usage in their outputs.
Progresses from basic filler removal to expert-level nested reasoning compression.
Contents
verbose_output - The padded, wasteful version of the text
efficient_output - The compressed, token-efficient equivalent
technique - Compression strategy used
subcategory - Specific variant of the technique
difficulty - Tier 1 (easiest)… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Token-Efficiency.Hindi_SFT_sentence_retriever_set
