CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01syafie-nzm /tokenized_datasettextn<1K0 likes7k downloads3y agoHugging Face02napaull /tokenized_C4textn<1K0 likes3.2k downloads5mo agoHugging Face03upup-ashton-wang /temp-decoder-train-tokenized1B<n<10B0 likes1.8k downloads5mo agoHugging Face04syafie-nzm /tokenized-dataset-combinetextn<1K0 likes1k downloads3y agoHugging Face05ks46 /urls-tokenized URLs (tokenized) ks46/urls-sampled run through a byte-level BPE built for URLs, stored as flat uint16 token streams that memory-map directly into a training loop. Shards 512 URLs 18,729,786,698 Tokens 664,731,047,208 Vocabulary 8,192 Token dtype uint16, little-endian There is no parquet here and the dataset viewer will not render it. These are raw token bins; see Reading the data below. Layout tokenizer/ the exact vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls-tokenized.tabulartext-generationn<1K0 likes620 downloads17d agoHugging Face06jinofy-corp /jora_corpus1_FR_tokenized_128ktabularn<1K2 likes614 downloads2mo agoHugging Face07qinglinhou /sokoban-10k-vjepa2-tokenizedtext10K<n<100K0 likes330 downloads5mo agoHugging Face08fillay /LWT-2.5B-model-6-tokenizedtabularn<1K0 likes301 downloads29d agoHugging Face090xBreath /holistic-tokenizedtext10K<n<100K0 likes295 downloads2y agoHugging Face10fillay /LWT-2.5B-model-5-tokenizedtabularn<1K0 likes218 downloads29d agoHugging Face11fillay /LWT-2.5B-model-7-tokenizedtabularn<1K0 likes216 downloads29d agoHugging Face12fillay /LWT-2.5B-model-4-tokenizedtabularn<1K0 likes196 downloads29d agoHugging Face13souvik18 /mistral_tokenized_2048_fixed_shards1M<n<10M0 likes192 downloads9mo agoHugging Face14fillay /LWT-2.5B-model-8-tokenizedtabularn<1K0 likes154 downloads29d agoHugging Face15fillay /LWT-2.5B-model-1-tokenizedtabularn<1K0 likes148 downloads29d agoHugging Face16fillay /LWT-2.5B-model-3-tokenizedtabularn<1K0 likes142 downloads29d agoHugging Face17fillay /LWT-2.5B-model-2-tokenizedtabularn<1K0 likes113 downloads29d agoHugging Face18fillay /LWT-2.5B-stage0-tokenizedtabularn<1K0 likes111 downloads1mo agoHugging Face19m-ric /TRM-modified-datamix-tokenized TRM modified datamix (tokenized) Pre-tokenized reasoning/pretraining mixture for from-scratch TRM (Tiny Recursive Model) training, built by running data_io — the HRM-Text data pipeline — verbatim on sapientinc/HRM-Text-data-io-cleaned-20260515, with three deliberate, documented deviations (below). It is emitted in the V1 tokenized dataset format (a single concatenated token pool + per-epoch document indices) and is ready to stream directly into training — no re-tokenization.… See the full description on the dataset page: https://huggingface.co/datasets/m-ric/TRM-modified-datamix-tokenized.tabulartext-generationn<1K0 likes58 downloads3mo agoHugging Face20instinct-org /miscellaneous_yt_chunked_tokenizedgated miscellaneous_yt_chunked_48k_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes39 downloads4mo agoHugging Face21noddu /swe_multiturn-claude-trace-tokenized-qwen3.5 SWE-bench Multi-turn Trace for LLMServingSim2 Derived from LMCache/Agentic-Traces, filtered to the claude-sonnet-4-6 + swebench subset, then re-tokenized with Qwen/Qwen3.5-122B-A10B to drive an LLM serving simulator with realistic per-turn prefix-cache reuse and strict turn-by-turn dependency enforcement. 110 sessions, 2,559 turns Average input 20.2K tokens per turn (max 65.9K) Average output 288.7 tokens per turn Per-session prefix-share ratio: 0.939 (turn N input is ~94% the same… See the full description on the dataset page: https://huggingface.co/datasets/noddu/swe_multiturn-claude-trace-tokenized-qwen3.5.tabular1K<n<10K0 likes36 downloads4mo agoHugging Face22gonglinyuan /code_search_net_python_tokenizedtext100K<n<1M2 likes34 downloads3y agoHugging Face23raghavnimbalkar /movie-screenplays-tokenized-dataset Screenplay Corpus — Tokenized (GPT-2) Pre-tokenized screenplay corpus used to train and evaluate the models in the GPT-2 Screenplay Fine-Tuning Study. Derived from Movie-Script-Database by Aveek Saha. Provided as tokenized JSON splits ready for direct consumption by a GPT-2 Trainer pipeline — no preprocessing required. Dataset Description This dataset contains approximately 94 million tokens of professionally formatted screenplay text, pre-tokenized using the… See the full description on the dataset page: https://huggingface.co/datasets/raghavnimbalkar/movie-screenplays-tokenized-dataset.text-generation100K<n<1M2 likes34 downloads4mo agoHugging Face24BootsofLagrangian /fineweb-edu-100b-smollmv2-tokenizedtextn<1K0 likes31 downloads9mo agoHugging Face25birgermoell /oellm-longctx-tokenized-superlong-512k-1m-2m-v1 OELLM Superlong Long-Context Tokenized 512K/1M/2M v1 This dataset is a superlong-context continuation-training add-on for extending beyond 256K toward 1M-2M context windows. The design goal is not simply longer packed text. At 1M-2M, the model needs collection-level continuity and explicit pressure to use very old evidence. This artifact therefore mixes real long structured/natural sources with a small multilingual full-span recall component. Source families: RFC Editor… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/oellm-longctx-tokenized-superlong-512k-1m-2m-v1.tabulartext-generationn<1K0 likes31 downloads3mo agoHugging Face26Kasmic /tokenized-crosscoder-qwen3-hhrf-dpo-pt Tokenized CrossCoder Dataset (Qwen3 + HH-RLHF + DPO) Description This dataset contains tokenized preference pairs from Anthropic's HH-RLHF and LMSys Chatbot Arena conversations, preprocessed for Direct Preference Optimization (DPO) training using the Qwen3-8B-Base tokenizer. Dataset Details Tokenizer: Qwen/Qwen3-8B-Base Max Sequence Length: 1024 tokens Format: PyTorch tensors (.pt file) Task: Direct Preference Optimization (DPO) Sources: Anthropic/hh-rlhf… See the full description on the dataset page: https://huggingface.co/datasets/Kasmic/tokenized-crosscoder-qwen3-hhrf-dpo-pt.textn<1K0 likes29 downloads1y agoHugging Face27Abby-Woodring /PEFT_tokenized_datasets100K<n<1M0 likes29 downloads1mo agoHugging Face28Kushala /tokenizedfiletextn<1K0 likes25 downloads2y agoHugging Face29BaiqingL /showdown-dojo-tokenized10K<n<100K0 likes24 downloads2y agoHugging Face30instinct-org /audiobook_chunked_tokenizedgated audiobook_chunked_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_tokenized.tabulartext-to-speech1M<n<10M0 likes20 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.