CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01syafie-nzm /tokenized_datasettextn<1K0 likes7k downloads3y agoHugging Face02TokenBender /code_instructions_122k_alpaca_styletext100K<n<1M80 likes5.8k downloads3y agoHugging Face03napaull /tokenized_C4textn<1K0 likes3.3k downloads5mo agoHugging Face04syafie-nzm /tokenized-dataset-combinetextn<1K0 likes821 downloads3y agoHugging Face05TokenWasteGroup /DynamicMCPBench DynamicMCPBench A trace-grounded, effect-scored benchmark for LLM agents on live MCP servers. Tasks are generated forward: an explorer agent drives real MCP tools until a goal is reached, the recorded trace is distilled into a TaskSpec, and candidates are graded on whether they reproduce the effects the trace produced — checkpoints, equivalence sets, minefields, a partial order — never on matching an answer string or a fixed tool list. Candidates are evaluated under… See the full description on the dataset page: https://huggingface.co/datasets/TokenWasteGroup/DynamicMCPBench.textother1K<n<10K1 likes739 downloads29d agoHugging Face06ks46 /urls-tokenized URLs (tokenized) ks46/urls-sampled run through a byte-level BPE built for URLs, stored as flat uint16 token streams that memory-map directly into a training loop. Shards 512 URLs 18,729,786,698 Tokens 664,731,047,208 Vocabulary 8,192 Token dtype uint16, little-endian There is no parquet here and the dataset viewer will not render it. These are raw token bins; see Reading the data below. Layout tokenizer/ the exact vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls-tokenized.tabulartext-generationn<1K0 likes622 downloads21d agoHugging Face07erickfmm /tokenised_subsetof_erickfmm__red_pajama_es_hq_35tabularn<1K0 likes592 downloads8mo agoHugging Face08Efe2898 /tokenizedtabularn<1K0 likes383 downloads3d agoHugging Face09MaxDevv /Qwen3.8-27B-Distill-1M-3.12B-Tokens Qwen3.8-27B-Distill-1M-4.83B-Tokens A unified, globally deduplicated, large-scale supervised distillation corpus built from 992,318 conversations generated by Qwen/Qwen3.8-27B, containing 4,834,771,862 target output tokens (3,570,459,498 reasoning tokens + 1,264,312,364 final response tokens) and 5,104,980,053 total sequence tokens. 1. Dataset Overview This dataset merges, aligns, and deduplicates the two primary high-quality Qwen3.8-27B generation corpora on… See the full description on the dataset page: https://huggingface.co/datasets/MaxDevv/Qwen3.8-27B-Distill-1M-3.12B-Tokens.texttext-generation100K<n<1M1 likes327 downloads28d agoHugging Face10qinglinhou /sokoban-10k-vjepa2-tokenizedtext10K<n<100K0 likes319 downloads6mo agoHugging Face11eduagarcia /multilingual_tokenizer_benchmark Multilingual Tokenizer Benchmark More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root. Natural language word count functions Download spacy models pip install ntlk spacy pygments underthesea camel-tools python -m spacy download ko_core_news_sm python -m spacy download ja_core_news_sm python -m spacy download zh_core_web_sm import nltk nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.tabulartext-generation100K<n<1M2 likes224 downloads1y agoHugging Face12Mr-Philo /dolma3_dolmino_megatron_tokenize Dolma 3 / Dolmino Megatron-LM indexed dataset This repository contains immutable Megatron-LM indexed datasets (.bin and .idx) produced from pinned Dolma 3 and Dolmino releases. It intentionally contains no training checkpoints, experiment outputs, logs, or dataset caches. The indexed payloads were derived from these pinned public datasets: allenai/dolma3_mix-150B-1025@afa92bfb22366821c5e6cd427cdd036b34b713ef… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Philo/dolma3_dolmino_megatron_tokenize.tabulartext-generationn<1K0 likes196 downloads27d agoHugging Face13jinofy-corp /jora_corpus1_FR_tokenized_128ktabularn<1K2 likes188 downloads2mo agoHugging Face14procmarco /fpabl1-arm-b-fp-tokens-48k fpabl1-arm-b-fp-tokens-48k Pre-tokenized bins for the FinePhrase vs FineWeb ablation (fpabl1), arm fpabl1-b-fp. Tokenizer: runs/mixed-tokenizer-48k/tokenizer (byte-BPE, 48k vocab; special ids bos=49119, eos=49120, pad=49121). Format: uint16 little-endian, 50 shards x 100,000,000 tokens = 5,000,000,000 tokens total. Boundary policy: source streams are already BOS/document/EOS packed; arm scheduler interleaves token chunks Source token composition: fp_en: 1,000,000,000 fp_ita:… See the full description on the dataset page: https://huggingface.co/datasets/procmarco/fpabl1-arm-b-fp-tokens-48k.tabularn<1K0 likes184 downloads3mo agoHugging Face15mr233 /TokenHD-eval-data TokenHD Evaluation Data Evaluation benchmarks for TokenHD, a pipeline for training token-level hallucination detectors in LLMs. Paper: arxiv.org/abs/2605.12384 Code: github.com/rmin2000/TokenHD Pre-trained Models: TokenHD Collection Training Data: mr233/TokenHD-training-data Benchmarks File benchmark value Domain SamplesIncorrect Correct tokenhd_eval_math_500.jsonl math_500 Math (MATH-500) 949 214 735 tokenhd_eval_math_aime.jsonl aime Math (AIME… See the full description on the dataset page: https://huggingface.co/datasets/mr233/TokenHD-eval-data.texttoken-classification1K<n<10K0 likes183 downloads5mo agoHugging Face16translorentz /vision-token-compression-bench OPTIC-Bench Optical Text In-Context Benchmark: how reliably do LLMs consume text delivered as rendered images versus plain text tokens? In summary, the evaluation reported here finds that optical text compression is effective only within a narrow and specific envelope. Delivering content as rendered images genuinely reduces input tokens, by thirteen to fifty-four per cent depending on the model and the language, but only when the document is long, the rendering is dense and the… See the full description on the dataset page: https://huggingface.co/datasets/translorentz/vision-token-compression-bench.imagevisual-question-answering1K<n<10K0 likes175 downloads3mo agoHugging Face17procesaur /sr-tokenizer-test Sr Tokenizer test This dataset provides a large Serbian text corpus designed for training and evaluating of tokenizers for Serbian language models. It combines multiple sources of Serbian text in both Cyrillic and Latin scripts, unified into a consistent JSONL format with id and text fields. Dataset Structure Metadata has been stripped; Each record is a JSON object with: id: unique identifier text: raw Serbian text Source coprora Znanje(sr) corpus: ~6.6 GB… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/sr-tokenizer-test.text100K<n<1M0 likes150 downloads5mo agoHugging Face18dougalldeepmind /2026-08-25-table2-9284-difficult-advice-verbose-token-matched-train-mixture Token-matched verbose difficult-advice arm. Holds difficult advice's share of the TRAINABLE TOKENS at the control's value while the traces are ~3x longer, by keeping only a subset of the expanded rows. Its sibling arm holds the ROW share instead; together they separate more deliberation from more difficult-advice signal. field value experiment Token-matched verbose difficult-advice arm. Holds difficult advice's share of the TRAINABLE TOKENS at the control's value… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-25-table2-9284-difficult-advice-verbose-token-matched-train-mixture.text1K<n<10K0 likes142 downloads1mo agoHugging Face19zjhhhh /fixed-n-rb-cost-aware-marginrl-qwen3-1.7b-base-math12k-token-mean-rerun-rollouts fixed_n_rb_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_token_mean_rerun rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes140 downloads21d agoHugging Face20Antix5 /brand-heavy-token-quality-datasettext1K<n<10K0 likes135 downloads1mo agoHugging Face21hamishivi /alpaca-farm-davinci-003-2048-tokentextn<1K3 likes134 downloads3y agoHugging Face22Antix5 /structure-heavy-token-quality-datasettext1K<n<10K0 likes112 downloads1mo agoHugging Face23TokenBender /python_evol_instruct_51ktext10K<n<100K6 likes105 downloads3y agoHugging Face24Antix5 /general-product-token-quality-datasettext1K<n<10K0 likes103 downloads1mo agoHugging Face25next-token /clean-PD-16000-books3 📚 clean-PD-16000-books3 A treasure trove of ~16,000 high-quality, public domain books in English language — nicely cleaned, with rich metadata, and ready for language modeling. ✨ What Makes This Dataset Special? This isn’t just another dump of dusty old text files. clean-PD-16000-books3 is the result of a rigorous cleaning and curation process applied to a large collection of public domain literature, including: ✅ Readable prose — paragraphized prose, without unnatural… See the full description on the dataset page: https://huggingface.co/datasets/next-token/clean-PD-16000-books3.text10K<n<100K6 likes100 downloads1y agoHugging Face26Pacific-i64 /data-32k-200b-tokens TR-HASH 32K · 200B Token Mixture Pretokenized training mixture for compact TR-HASH language-model research. The Dataset Viewer displays one summary row per source. The actual training data is stored as packed uint16 token shards under corpora/<source>/tokens-*.bin. Each sequence contains 1,024 token IDs produced by the project 32K tokenizer. Mixture Source Weight Training tokens DCLM 45% 90B FineWeb-Edu deduplicated 30% 60B Stack-Edu 10% 20B… See the full description on the dataset page: https://huggingface.co/datasets/Pacific-i64/data-32k-200b-tokens.tabulartext-generationn<1K1 likes99 downloads1mo agoHugging Face27pcuenq /tokenizer-conformance Tokenizer conformance fixtures Reference inputs and Python fast-tokenizer outputs for tokenizer implementations. The initial corpus contains 83 inputs in 30 categories, with 498 reference encodings across six tokenizers. This is a regression dataset, not a model-quality benchmark. Provenance and attribution The input corpus and reference entries come from apocryphx's swift-transformers PR #360, at commit ce847085784bacd8c3c15180c976b17c8ce73e31. The corpus… See the full description on the dataset page: https://huggingface.co/datasets/pcuenq/tokenizer-conformance.textothern<1K0 likes94 downloads6d agoHugging Face28Antix5 /food-product-token-quality-datasettext10K<n<100K0 likes90 downloads1mo agoHugging Face29zjhhhh /fixed-n-rb-er-cost-marginrl-qwen3-1.7b-base-math12k-token-mean-run2-rollouts fixed_n_rb_er_cost_marginrl_Qwen3-1.7B-Base_math12k_token_mean_run2 rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes85 downloads20d agoHugging Face30kacperwikiel /speakleash-tokenizer-5gb-sample SpeakLeash tokenizer 42GB quality sample Private tokenizer-training sample built from a stratified local SpeakLeash text snapshot on gb10. Target: 5.0 GiB raw JSONL UTF-8 bytes after document filters and exact text dedup. Format: data/tokenizer_sample_*.jsonl.zst, one JSON object per line with text, source, source_kind. This is intended for tokenizer/BPE training convergence tests. texttext-generation1M<n<10M0 likes75 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.