datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tokenized_datasetcode_instructions_122k_alpaca_styleEmilia-Dataset-tokenisedtokenized_C4temp-decoder-train-tokenizedtokenized-dataset-combinetokenised_subsetof_erickfmm__red_pajama_es_hq_35DynamicMCPBench
DynamicMCPBench
A trace-grounded, effect-scored benchmark for LLM agents on live MCP servers.
Tasks are generated forward: an explorer agent drives real MCP tools until a goal
is reached, the recorded trace is distilled into a TaskSpec, and candidates are graded
on whether they reproduce the effects the trace produced — checkpoints, equivalence
sets, minefields, a partial order — never on matching an answer string or a fixed tool
list. Candidates are evaluated under… See the full description on the dataset page: https://huggingface.co/datasets/TokenWasteGroup/DynamicMCPBench.urls-tokenized
URLs (tokenized)
ks46/urls-sampled run through a byte-level
BPE built for URLs, stored as flat uint16 token streams that memory-map
directly into a training loop.
Shards
512
URLs
18,729,786,698
Tokens
664,731,047,208
Vocabulary
8,192
Token dtype
uint16, little-endian
There is no parquet here and the dataset viewer will not render it. These
are raw token bins; see Reading the data below.
Layout
tokenizer/ the exact vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls-tokenized.jora_corpus1_FR_tokenized_128kbrand-heavy-token-quality-datasetfood-product-token-quality-datasetgeneral-product-token-quality-dataset2026-08-25-table2-9284-difficult-advice-verbose-token-matched-train-mixture
Token-matched verbose difficult-advice arm. Holds difficult advice's share of the TRAINABLE TOKENS at the control's value while the traces are ~3x longer, by keeping only a subset of the expanded rows. Its sibling arm holds the ROW share instead; together they separate more deliberation from more difficult-advice signal.
field
value
experiment
Token-matched verbose difficult-advice arm. Holds difficult advice's share of the TRAINABLE TOKENS at the control's value… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-25-table2-9284-difficult-advice-verbose-token-matched-train-mixture.sokoban-10k-vjepa2-tokenizedstructure-heavy-token-quality-datasetfixed-n-rb-er-cost-marginrl-qwen3-1.7b-base-math12k-token-mean-fixed-q0p8-run2-rollouts
fixed_n_rb_er_cost_marginrl_Qwen3-1.7B-Base_math12k_token_mean_fixed_q0.8_run2 rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
LWT-2.5B-model-6-tokenizedQwen3.8-27B-Distill-1M-3.12B-Tokens
Qwen3.8-27B-Distill-1M-4.83B-Tokens
A unified, globally deduplicated, large-scale supervised distillation corpus built from 992,318 conversations generated by Qwen/Qwen3.8-27B, containing 4,834,771,862 target output tokens (3,570,459,498 reasoning tokens + 1,264,312,364 final response tokens) and 5,104,980,053 total sequence tokens.
1. Dataset Overview
This dataset merges, aligns, and deduplicates the two primary high-quality Qwen3.8-27B generation corpora on… See the full description on the dataset page: https://huggingface.co/datasets/MaxDevv/Qwen3.8-27B-Distill-1M-3.12B-Tokens.holistic-tokenizedmultilingual_tokenizer_benchmark
Multilingual Tokenizer Benchmark
More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root.
Natural language word count functions
Download spacy models
pip install ntlk spacy pygments underthesea camel-tools
python -m spacy download ko_core_news_sm
python -m spacy download ja_core_news_sm
python -m spacy download zh_core_web_sm
import nltk
nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.metrics-outputs-pcfg-matryoshka-layer-03-token-cacheLWT-2.5B-model-5-tokenizedLWT-2.5B-model-7-tokenizedmetrics-outputs-pcfg-matryoshka-layer-02-token-cacheLWT-2.5B-model-4-tokenizedmistral_tokenized_2048_fixed_shardsdolma3_dolmino_megatron_tokenize
Dolma 3 / Dolmino Megatron-LM indexed dataset
This repository contains immutable Megatron-LM indexed datasets (.bin and
.idx) produced from pinned Dolma 3 and Dolmino releases. It intentionally
contains no training checkpoints, experiment outputs, logs, or dataset caches.
The indexed payloads were derived from these pinned public datasets:
allenai/dolma3_mix-150B-1025@afa92bfb22366821c5e6cd427cdd036b34b713ef… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Philo/dolma3_dolmino_megatron_tokenize.data-32k-200b-tokens
TR-HASH 32K · 200B Token Mixture
Pretokenized training mixture for compact TR-HASH language-model research.
The Dataset Viewer displays one summary row per source. The actual training data
is stored as packed uint16 token shards under corpora/<source>/tokens-*.bin.
Each sequence contains 1,024 token IDs produced by the project 32K tokenizer.
Mixture
Source
Weight
Training tokens
DCLM
45%
90B
FineWeb-Edu deduplicated
30%
60B
Stack-Edu
10%
20B… See the full description on the dataset page: https://huggingface.co/datasets/AETHORIA-AI/data-32k-200b-tokens.metrics-outputs-pcfg-matryoshka-layer-01-token-cache
