datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tokenized_datasetcode_instructions_122k_alpaca_styletokenized_C4tokenized-dataset-combineDynamicMCPBench
DynamicMCPBench
A trace-grounded, effect-scored benchmark for LLM agents on live MCP servers.
Tasks are generated forward: an explorer agent drives real MCP tools until a goal
is reached, the recorded trace is distilled into a TaskSpec, and candidates are graded
on whether they reproduce the effects the trace produced — checkpoints, equivalence
sets, minefields, a partial order — never on matching an answer string or a fixed tool
list. Candidates are evaluated under… See the full description on the dataset page: https://huggingface.co/datasets/TokenWasteGroup/DynamicMCPBench.urls-tokenized
URLs (tokenized)
ks46/urls-sampled run through a byte-level
BPE built for URLs, stored as flat uint16 token streams that memory-map
directly into a training loop.
Shards
512
URLs
18,729,786,698
Tokens
664,731,047,208
Vocabulary
8,192
Token dtype
uint16, little-endian
There is no parquet here and the dataset viewer will not render it. These
are raw token bins; see Reading the data below.
Layout
tokenizer/ the exact vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls-tokenized.tokenised_subsetof_erickfmm__red_pajama_es_hq_35tokenizedQwen3.8-27B-Distill-1M-3.12B-Tokens
Qwen3.8-27B-Distill-1M-4.83B-Tokens
A unified, globally deduplicated, large-scale supervised distillation corpus built from 992,318 conversations generated by Qwen/Qwen3.8-27B, containing 4,834,771,862 target output tokens (3,570,459,498 reasoning tokens + 1,264,312,364 final response tokens) and 5,104,980,053 total sequence tokens.
1. Dataset Overview
This dataset merges, aligns, and deduplicates the two primary high-quality Qwen3.8-27B generation corpora on… See the full description on the dataset page: https://huggingface.co/datasets/MaxDevv/Qwen3.8-27B-Distill-1M-3.12B-Tokens.sokoban-10k-vjepa2-tokenizedmultilingual_tokenizer_benchmark
Multilingual Tokenizer Benchmark
More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root.
Natural language word count functions
Download spacy models
pip install ntlk spacy pygments underthesea camel-tools
python -m spacy download ko_core_news_sm
python -m spacy download ja_core_news_sm
python -m spacy download zh_core_web_sm
import nltk
nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.dolma3_dolmino_megatron_tokenize
Dolma 3 / Dolmino Megatron-LM indexed dataset
This repository contains immutable Megatron-LM indexed datasets (.bin and
.idx) produced from pinned Dolma 3 and Dolmino releases. It intentionally
contains no training checkpoints, experiment outputs, logs, or dataset caches.
The indexed payloads were derived from these pinned public datasets:
allenai/dolma3_mix-150B-1025@afa92bfb22366821c5e6cd427cdd036b34b713ef… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Philo/dolma3_dolmino_megatron_tokenize.jora_corpus1_FR_tokenized_128kfpabl1-arm-b-fp-tokens-48k
fpabl1-arm-b-fp-tokens-48k
Pre-tokenized bins for the FinePhrase vs FineWeb ablation (fpabl1), arm fpabl1-b-fp.
Tokenizer: runs/mixed-tokenizer-48k/tokenizer (byte-BPE, 48k vocab; special ids bos=49119, eos=49120, pad=49121).
Format: uint16 little-endian, 50 shards x 100,000,000 tokens = 5,000,000,000 tokens total.
Boundary policy: source streams are already BOS/document/EOS packed; arm scheduler interleaves token chunks
Source token composition:
fp_en: 1,000,000,000
fp_ita:… See the full description on the dataset page: https://huggingface.co/datasets/procmarco/fpabl1-arm-b-fp-tokens-48k.TokenHD-eval-data
TokenHD Evaluation Data
Evaluation benchmarks for TokenHD, a pipeline for training token-level hallucination detectors in LLMs.
Paper: arxiv.org/abs/2605.12384
Code: github.com/rmin2000/TokenHD
Pre-trained Models: TokenHD Collection
Training Data: mr233/TokenHD-training-data
Benchmarks
File
benchmark value
Domain
SamplesIncorrect
Correct
tokenhd_eval_math_500.jsonl
math_500
Math (MATH-500)
949
214
735
tokenhd_eval_math_aime.jsonl
aime
Math (AIME… See the full description on the dataset page: https://huggingface.co/datasets/mr233/TokenHD-eval-data.vision-token-compression-bench
OPTIC-Bench
Optical Text In-Context Benchmark: how reliably do LLMs consume text
delivered as rendered images versus plain text tokens?
In summary, the evaluation reported here finds that optical text compression
is effective only within a narrow and specific envelope. Delivering content
as rendered images genuinely reduces input tokens, by thirteen to
fifty-four per cent depending on the model and the language, but only when
the document is long, the rendering is dense and the… See the full description on the dataset page: https://huggingface.co/datasets/translorentz/vision-token-compression-bench.sr-tokenizer-test
Sr Tokenizer test
This dataset provides a large Serbian text corpus designed for training and evaluating of tokenizers for Serbian language models.
It combines multiple sources of Serbian text in both Cyrillic and Latin scripts, unified into a consistent JSONL format with id and text fields.
Dataset Structure
Metadata has been stripped; Each record is a JSON object with:
id: unique identifier
text: raw Serbian text
Source coprora
Znanje(sr) corpus: ~6.6 GB… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/sr-tokenizer-test.2026-08-25-table2-9284-difficult-advice-verbose-token-matched-train-mixture
Token-matched verbose difficult-advice arm. Holds difficult advice's share of the TRAINABLE TOKENS at the control's value while the traces are ~3x longer, by keeping only a subset of the expanded rows. Its sibling arm holds the ROW share instead; together they separate more deliberation from more difficult-advice signal.
field
value
experiment
Token-matched verbose difficult-advice arm. Holds difficult advice's share of the TRAINABLE TOKENS at the control's value… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-25-table2-9284-difficult-advice-verbose-token-matched-train-mixture.fixed-n-rb-cost-aware-marginrl-qwen3-1.7b-base-math12k-token-mean-rerun-rollouts
fixed_n_rb_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_token_mean_rerun rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
brand-heavy-token-quality-datasetalpaca-farm-davinci-003-2048-tokenstructure-heavy-token-quality-datasetpython_evol_instruct_51kgeneral-product-token-quality-datasetclean-PD-16000-books3
📚 clean-PD-16000-books3
A treasure trove of ~16,000 high-quality, public domain books in English language — nicely cleaned, with rich metadata, and ready for language modeling.
✨ What Makes This Dataset Special?
This isn’t just another dump of dusty old text files.
clean-PD-16000-books3 is the result of a rigorous cleaning and curation process applied to a large collection of public domain literature, including:
✅ Readable prose — paragraphized prose, without unnatural… See the full description on the dataset page: https://huggingface.co/datasets/next-token/clean-PD-16000-books3.data-32k-200b-tokens
TR-HASH 32K · 200B Token Mixture
Pretokenized training mixture for compact TR-HASH language-model research.
The Dataset Viewer displays one summary row per source. The actual training data
is stored as packed uint16 token shards under corpora/<source>/tokens-*.bin.
Each sequence contains 1,024 token IDs produced by the project 32K tokenizer.
Mixture
Source
Weight
Training tokens
DCLM
45%
90B
FineWeb-Edu deduplicated
30%
60B
Stack-Edu
10%
20B… See the full description on the dataset page: https://huggingface.co/datasets/Pacific-i64/data-32k-200b-tokens.tokenizer-conformance
Tokenizer conformance fixtures
Reference inputs and Python fast-tokenizer outputs for tokenizer implementations.
The initial corpus contains 83 inputs in 30 categories, with 498 reference encodings across six tokenizers.
This is a regression dataset, not a model-quality benchmark.
Provenance and attribution
The input corpus and reference entries come from apocryphx's swift-transformers PR #360,
at commit ce847085784bacd8c3c15180c976b17c8ce73e31.
The corpus… See the full description on the dataset page: https://huggingface.co/datasets/pcuenq/tokenizer-conformance.food-product-token-quality-datasetfixed-n-rb-er-cost-marginrl-qwen3-1.7b-base-math12k-token-mean-run2-rollouts
fixed_n_rb_er_cost_marginrl_Qwen3-1.7B-Base_math12k_token_mean_run2 rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
speakleash-tokenizer-5gb-sample
SpeakLeash tokenizer 42GB quality sample
Private tokenizer-training sample built from a stratified local SpeakLeash text snapshot on gb10.
Target: 5.0 GiB raw JSONL UTF-8 bytes after document filters and exact text dedup.
Format: data/tokenizer_sample_*.jsonl.zst, one JSON object per line with text, source, source_kind.
This is intended for tokenizer/BPE training convergence tests.
