datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
token-counts
Marin Token Counts
Token counts for all datasets used in Marin pretraining runs.
Schema
Column
Type
Description
dataset
string
Dataset identifier
marin_tokens
int
Number of tokens after tokenization
category
string
Content domain (web, code, math, academic, books, etc.)
synthetic
bool
Whether the data is LLM-generated or LLM-translated
Categories
web — Quality-classified Common Crawl text (Nemotron-CC)
code — Source code and… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/token-counts.parler-tts_mls_eng_10k_snac_token_old
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.fixed-tokenizer-morphscore-segmentsablation_tokensprompts_under_512_tokens
Under 512 Tokens Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use
Dataset Overview
Specialized collection of short-form English prompts (under 512 tokens), perfect for training models with context length constraints or faster iteration cycles.
📊 Dataset Statistics
Metric
Value
Total Files
200
Rows Per File
10,000
Total Rows
2,000,000
Token Range
1 to 511 tokens… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/prompts_under_512_tokens.globalise_NER_token_classification_dataset
Dataset Card for Dataset Name
The globalise_NER_token_classification dataset is a fine-grained dataset for the training of token-classification NER models on Dutch East-India Company texts (17th to 18th century).
Dataset Details
Dataset Description
The dataset provides 15 fine-grained labels detailing activities and people of the Dutch East-India Company (VOC), and can be used to train NER token-classification models for the
period 17th-18th century and the… See the full description on the dataset page: https://huggingface.co/datasets/globalise/globalise_NER_token_classification_dataset.Piyyuttokenizer-leaderboard
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): en
License: mit
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Direct Use
[More… See the full description on the dataset page: https://huggingface.co/datasets/Lyte/tokenizer-leaderboard.classification_token_propagandaempathetic_dialogues_with_special_tokenstokopedia-product-reviews-2019
Tokopedia Product Reviews 2019
Dataset Description
This dataset contains 40,607 product reviews from Tokopedia, one of Indonesia's largest e-commerce platforms, scraped in 2019. The dataset provides valuable insights into customer sentiment and shopping behavior in the Indonesian e-commerce market.
Dataset Summary
Language: Indonesian (Bahasa Indonesia)
Task: Sentiment Analysis, Product Review Analysis, E-commerce Research
Size: 40,607 reviews
Categories: 5… See the full description on the dataset page: https://huggingface.co/datasets/farhamu/tokopedia-product-reviews-2019.JBB-Behaviors
An Open Robustness Benchmark for Jailbreaking Language Models
NeurIPS 2024 Datasets and Benchmarks Track
Paper |
Leaderboard |
Benchmark code
What is JailbreakBench?
Jailbreakbench is an open-source robustness benchmark for jailbreaking large language models (LLMs). The goal of this benchmark is to comprehensively track progress toward (1) generating successful jailbreaks and (2) defending against these jailbreaks. To this end, we… See the full description on the dataset page: https://huggingface.co/datasets/Tokyomonster/JBB-Behaviors.webnlg_tokensmedium_512_1k_tokens_prompts
Medium 512-1K Tokens Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ By using this dataset you agree to our Terms of Use.
Overview
703 high-quality English prompts whose length lies between 512 and 1 000 tokens.Every prompt has been de-duplicated, cleaned and token-counted with the GPT-2 tokenizer.
Statistics
Rows
Token range
File size
Format
703
512 – 1 000
2.9 MB
CSV
Use-cases
Medium-context language-model fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/medium_512_1k_tokens_prompts.MMLU-Pro-single-token-entropy
Dataset Card for MMLU Pro with single token response entropy metadata for Mistral 24B, Phi4, Phi4-mini, Qwen2.5 3B
MMLU Pro dataset with single token response entropy metadata for Mistral 24B, Phi4, Phi4-mini, Qwen2.5 3B
Dataset Details
Dataset Description
Following up on the results from "When an LLM is apprehensive about its answers -- and when its uncertainty is justified", we measure the response entopy for MMLU Pro dataset when the model is prompted to… See the full description on the dataset page: https://huggingface.co/datasets/LabARSS/MMLU-Pro-single-token-entropy.long_over_1k_tokens_prompts
Long Over 1K Tokens Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use
Dataset Overview
Specialized collection of long-form English prompts (≥ 1 000 tokens) for training advanced models that require extensive context and complex reasoning.
📊 Dataset Statistics
Metric
Value
Total Rows
289
Token Range
1 001 – 10 000 tokens
File Size
≈ 3.7 MB
Format
Single CSV file
Target… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/long_over_1k_tokens_prompts.SimpleDC
simpledc-dataset
Official huggingface dataset for the SimpleDC (Simple Digestive Cancer) dataset
Please cite as:
@article{rahman2024health,
title={Health Text Simplification: An Annotated Corpus for Digestive Cancer Education and Novel Strategies for Reinforcement Learning},
author={Rahman, Md Mushfiqur and Irbaz, Mohammad Sabik and North, Kai and Williams, Michelle S and Zampieri, Marcos and Lybarger, Kevin},
journal={arXiv preprint arXiv:2401.15043},
year={2024}
}
PubChem10M_SELFIES_TokenizedCustom cl100k tokenized version of PubChem10M_SELFIES.
whisper_transcriptions_token_idsDynamic-Topic-RedPajama-Data-1T-100k-SubSample-max-1k-tokens
Dynamic Topic Modeling Dataset: RedPajama-1T SubSample (100k samples, 1k tokens)
📝Check out the Blog Post
This dataset represents a curated subset of the RedPajama-1T Sample dataset, specifically processed for dynamic topic modeling applications. It contains 100,000
samples from the original dataset, with each document limited to the first 1,024 tokens for consistent processing.
Dataset Overview
Name:… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/Dynamic-Topic-RedPajama-Data-1T-100k-SubSample-max-1k-tokens.Token_Optimization_Org
AI Safety & Bias Evaluation Conversations
Dataset Summary
This dataset contains simulated multi-turn conversations designed to evaluate AI language model behavior across two safety-critical domains: self-harm response handling and political bias. Each row represents a single evaluation scenario where an AI model's responses are assessed for safety compliance or neutrality. The dataset is intended to support research and development of safer, less biased AI systems.
All… See the full description on the dataset page: https://huggingface.co/datasets/token-opt-org/Token_Optimization_Org.token-budgets-catalog
Token Budgets — Empirical catalogue and inter-rater reliability data
Data for:
Token Budgets: An Empirical Catalog of 63 LLM-Agent Budget-Overrun
Incidents, with an Affine-Typed Rust Mitigation as a Case Study.
Sajjad Khan, 2026.
arXiv:2606.04056 — preprint.
This dataset bundles three artefacts referenced in the paper:
catalogue — the harvested catalogue of LLM-agent budget-overrun
incidents across 21 orchestration frameworks (2023–2026), 167 rows
total. The IRR-included… See the full description on the dataset page: https://huggingface.co/datasets/sajjadanwar0/token-budgets-catalog.tokyo-vpn-monitor
language:
ja
en
license: mit
multilinguality:
multilingual
size_categories:
1K<n<10K
source_datasets:
original
task_categories:
other
task_ids: []
pretty_name: Tokyo VPN Speed Monitor Dataset
tags:
vpn
network-monitoring
performance-measurement
time-series
networking
internet-measurement
automated-testing
zero-cost-infrastructure
google-apps-script
Tokyo VPN Speed Monitor Dataset
Dataset Summary
The Tokyo VPN Speed Monitor Dataset contains continuous automated… See the full description on the dataset page: https://huggingface.co/datasets/blstweb0901/tokyo-vpn-monitor.tokenized-IELTS-writing-task-2-evaluation-DialoGPT-mediumBatch_indexing_machine_tokensrams-no-special-tokenssentence_retrieval_hindi_SFTparis-vs-tokyo-hotels-2026
Paris vs Tokyo Hotels 2026: Stars & Guest Ratings
Hotels in Paris and Tokyo (3 to 5 star), each with hotel name, platform star level and a blended guest rating (Booking, Priceline, Agoda, HotelsCombined). Collected 2026-06-22 by MrBridge.
Files
kayak_paris_tokyo_2026.csv — 197 hotels (Paris 98, Tokyo 99), the primary blended-OTA source.
priceline_paris_tokyo_2026.csv — 62 hotels, an independent Priceline pull used as a robustness check.
Columns: city, name… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Bridge/paris-vs-tokyo-hotels-2026.token-optimization
AI Safety & Bias Evaluation Conversations
Dataset Summary
This dataset contains simulated multi-turn conversations designed to evaluate AI language model behavior across two safety-critical domains: self-harm response handling and political bias. Each row represents a single evaluation scenario where an AI model's responses are assessed for safety compliance or neutrality. The dataset is intended to support research and development of safer, less biased AI… See the full description on the dataset page: https://huggingface.co/datasets/CentificAIResearch/token-optimization.Token-Efficiency
token_efficiency_corpus
A 2.5 GB CSV corpus teaching LLMs to minimize token usage in their outputs.
Progresses from basic filler removal to expert-level nested reasoning compression.
Contents
verbose_output - The padded, wasteful version of the text
efficient_output - The compressed, token-efficient equivalent
technique - Compression strategy used
subcategory - Specific variant of the technique
difficulty - Tier 1 (easiest)… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Token-Efficiency.
