CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01syafie-nzm /tokenized_datasettextn<1K0 likes7k downloads3y agoHugging Face02TokenBender /code_instructions_122k_alpaca_styletext100K<n<1M80 likes5.4k downloads3y agoHugging Face03taozi555 /Emilia-Dataset-tokenised10M<n<100M0 likes3.6k downloads1y agoHugging Face04napaull /tokenized_C4textn<1K0 likes3.2k downloads5mo agoHugging Face05upup-ashton-wang /temp-decoder-train-tokenized1B<n<10B0 likes1.9k downloads5mo agoHugging Face06syafie-nzm /tokenized-dataset-combinetextn<1K0 likes1k downloads3y agoHugging Face07erickfmm /tokenised_subsetof_erickfmm__red_pajama_es_hq_35tabularn<1K0 likes845 downloads7mo agoHugging Face08TokenWasteGroup /DynamicMCPBench DynamicMCPBench A trace-grounded, effect-scored benchmark for LLM agents on live MCP servers. Tasks are generated forward: an explorer agent drives real MCP tools until a goal is reached, the recorded trace is distilled into a TaskSpec, and candidates are graded on whether they reproduce the effects the trace produced — checkpoints, equivalence sets, minefields, a partial order — never on matching an answer string or a fixed tool list. Candidates are evaluated under… See the full description on the dataset page: https://huggingface.co/datasets/TokenWasteGroup/DynamicMCPBench.textother1K<n<10K1 likes648 downloads25d agoHugging Face09ks46 /urls-tokenized URLs (tokenized) ks46/urls-sampled run through a byte-level BPE built for URLs, stored as flat uint16 token streams that memory-map directly into a training loop. Shards 512 URLs 18,729,786,698 Tokens 664,731,047,208 Vocabulary 8,192 Token dtype uint16, little-endian There is no parquet here and the dataset viewer will not render it. These are raw token bins; see Reading the data below. Layout tokenizer/ the exact vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls-tokenized.tabulartext-generationn<1K0 likes621 downloads17d agoHugging Face10jinofy-corp /jora_corpus1_FR_tokenized_128ktabularn<1K2 likes614 downloads2mo agoHugging Face11Antix5 /brand-heavy-token-quality-datasettext1K<n<10K0 likes345 downloads28d agoHugging Face12Antix5 /food-product-token-quality-datasettext10K<n<100K0 likes344 downloads29d agoHugging Face13Antix5 /general-product-token-quality-datasettext1K<n<10K0 likes342 downloads29d agoHugging Face14dougalldeepmind /2026-08-25-table2-9284-difficult-advice-verbose-token-matched-train-mixture Token-matched verbose difficult-advice arm. Holds difficult advice's share of the TRAINABLE TOKENS at the control's value while the traces are ~3x longer, by keeping only a subset of the expanded rows. Its sibling arm holds the ROW share instead; together they separate more deliberation from more difficult-advice signal. field value experiment Token-matched verbose difficult-advice arm. Holds difficult advice's share of the TRAINABLE TOKENS at the control's value… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-25-table2-9284-difficult-advice-verbose-token-matched-train-mixture.text1K<n<10K0 likes341 downloads28d agoHugging Face15qinglinhou /sokoban-10k-vjepa2-tokenizedtext10K<n<100K0 likes325 downloads5mo agoHugging Face16Antix5 /structure-heavy-token-quality-datasettext1K<n<10K0 likes315 downloads28d agoHugging Face17zjhhhh /fixed-n-rb-er-cost-marginrl-qwen3-1.7b-base-math12k-token-mean-fixed-q0p8-run2-rollouts fixed_n_rb_er_cost_marginrl_Qwen3-1.7B-Base_math12k_token_mean_fixed_q0.8_run2 rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes311 downloads29d agoHugging Face18fillay /LWT-2.5B-model-6-tokenizedtabularn<1K0 likes301 downloads29d agoHugging Face19MaxDevv /Qwen3.8-27B-Distill-1M-3.12B-Tokens Qwen3.8-27B-Distill-1M-4.83B-Tokens A unified, globally deduplicated, large-scale supervised distillation corpus built from 992,318 conversations generated by Qwen/Qwen3.8-27B, containing 4,834,771,862 target output tokens (3,570,459,498 reasoning tokens + 1,264,312,364 final response tokens) and 5,104,980,053 total sequence tokens. 1. Dataset Overview This dataset merges, aligns, and deduplicates the two primary high-quality Qwen3.8-27B generation corpora on… See the full description on the dataset page: https://huggingface.co/datasets/MaxDevv/Qwen3.8-27B-Distill-1M-3.12B-Tokens.texttext-generation100K<n<1M1 likes290 downloads24d agoHugging Face200xBreath /holistic-tokenizedtext10K<n<100K0 likes268 downloads2y agoHugging Face21eduagarcia /multilingual_tokenizer_benchmark Multilingual Tokenizer Benchmark More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root. Natural language word count functions Download spacy models pip install ntlk spacy pygments underthesea camel-tools python -m spacy download ko_core_news_sm python -m spacy download ja_core_news_sm python -m spacy download zh_core_web_sm import nltk nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.tabulartext-generation100K<n<1M2 likes219 downloads1y agoHugging Face22soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-layer-03-token-cachetabularn<1K0 likes218 downloads1mo agoHugging Face23fillay /LWT-2.5B-model-5-tokenizedtabularn<1K0 likes218 downloads29d agoHugging Face24fillay /LWT-2.5B-model-7-tokenizedtabularn<1K0 likes216 downloads29d agoHugging Face25soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-layer-02-token-cachetabularn<1K0 likes197 downloads1mo agoHugging Face26fillay /LWT-2.5B-model-4-tokenizedtabularn<1K0 likes196 downloads29d agoHugging Face27souvik18 /mistral_tokenized_2048_fixed_shards1M<n<10M0 likes192 downloads9mo agoHugging Face28Mr-Philo /dolma3_dolmino_megatron_tokenize Dolma 3 / Dolmino Megatron-LM indexed dataset This repository contains immutable Megatron-LM indexed datasets (.bin and .idx) produced from pinned Dolma 3 and Dolmino releases. It intentionally contains no training checkpoints, experiment outputs, logs, or dataset caches. The indexed payloads were derived from these pinned public datasets: allenai/dolma3_mix-150B-1025@afa92bfb22366821c5e6cd427cdd036b34b713ef… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Philo/dolma3_dolmino_megatron_tokenize.tabulartext-generationn<1K0 likes191 downloads22d agoHugging Face29AETHORIA-AI /data-32k-200b-tokens TR-HASH 32K · 200B Token Mixture Pretokenized training mixture for compact TR-HASH language-model research. The Dataset Viewer displays one summary row per source. The actual training data is stored as packed uint16 token shards under corpora/<source>/tokens-*.bin. Each sequence contains 1,024 token IDs produced by the project 32K tokenizer. Mixture Source Weight Training tokens DCLM 45% 90B FineWeb-Edu deduplicated 30% 60B Stack-Edu 10% 20B… See the full description on the dataset page: https://huggingface.co/datasets/AETHORIA-AI/data-32k-200b-tokens.tabulartext-generationn<1K0 likes182 downloads1mo agoHugging Face30soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-layer-01-token-cachetabularn<1K0 likes181 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.