CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ks46 /urls-tokenized URLs (tokenized) ks46/urls-sampled run through a byte-level BPE built for URLs, stored as flat uint16 token streams that memory-map directly into a training loop. Shards 512 URLs 18,729,786,698 Tokens 664,731,047,208 Vocabulary 8,192 Token dtype uint16, little-endian There is no parquet here and the dataset viewer will not render it. These are raw token bins; see Reading the data below. Layout tokenizer/ the exact vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls-tokenized.tabulartext-generationn<1K0 likes621 downloads18d agoHugging Face02m-ric /TRM-modified-datamix-tokenized TRM modified datamix (tokenized) Pre-tokenized reasoning/pretraining mixture for from-scratch TRM (Tiny Recursive Model) training, built by running data_io — the HRM-Text data pipeline — verbatim on sapientinc/HRM-Text-data-io-cleaned-20260515, with three deliberate, documented deviations (below). It is emitted in the V1 tokenized dataset format (a single concatenated token pool + per-epoch document indices) and is ready to stream directly into training — no re-tokenization.… See the full description on the dataset page: https://huggingface.co/datasets/m-ric/TRM-modified-datamix-tokenized.tabulartext-generationn<1K0 likes51 downloads3mo agoHugging Face03raghavnimbalkar /movie-screenplays-tokenized-dataset Screenplay Corpus — Tokenized (GPT-2) Pre-tokenized screenplay corpus used to train and evaluate the models in the GPT-2 Screenplay Fine-Tuning Study. Derived from Movie-Script-Database by Aveek Saha. Provided as tokenized JSON splits ready for direct consumption by a GPT-2 Trainer pipeline — no preprocessing required. Dataset Description This dataset contains approximately 94 million tokens of professionally formatted screenplay text, pre-tokenized using the… See the full description on the dataset page: https://huggingface.co/datasets/raghavnimbalkar/movie-screenplays-tokenized-dataset.text-generation100K<n<1M2 likes49 downloads4mo agoHugging Face04birgermoell /oellm-longctx-tokenized-superlong-512k-1m-2m-v1 OELLM Superlong Long-Context Tokenized 512K/1M/2M v1 This dataset is a superlong-context continuation-training add-on for extending beyond 256K toward 1M-2M context windows. The design goal is not simply longer packed text. At 1M-2M, the model needs collection-level continuity and explicit pressure to use very old evidence. This artifact therefore mixes real long structured/natural sources with a small multilingual full-span recall component. Source families: RFC Editor… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/oellm-longctx-tokenized-superlong-512k-1m-2m-v1.tabulartext-generationn<1K0 likes32 downloads3mo agoHugging Face05EmpathicRobotics /FineVideo-Prototype-Tokenized FineVideo-Prototype-Tokenized — Base Video Token Dataset Overview This dataset contains the base video tokenization output from the prototype pipeline, extracted from ~40K YouTube videos in the FineVideo dataset. Each video is tokenised into three modalities: Seed2 — 1 FPS semantic keyframe tokens (vocab: 8,192) Cosmos — every 8 frames spatial video tokens (vocab: 64,000) AVC-LM — every 8 frames H.264 BPE tokens (vocab: 8,192) This dataset does not contain 3D… See the full description on the dataset page: https://huggingface.co/datasets/EmpathicRobotics/FineVideo-Prototype-Tokenized.textvideo-classification10K<n<100K0 likes16 downloads3mo agoHugging Face06agentlans /c4-en-tokenized C4 English Tokenized Samples This dataset contains tokenized English samples from the C4 (Colossal Clean Crawled Corpus) dataset for natural language processing (NLP) tasks. The first 125 000 entries from the en split of allenai/c4 were tokenized using spaCy's en_core_web_sm model. Tokens joined with spaces. Features text: Original text from C4 tokenized: The tokenized and space-joined text num_tokens: Number of tokens after tokenization num_punct_tokens: Number of… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/c4-en-tokenized.tabulartext-generation100K<n<1M0 likes8 downloads2y agoHugging Face07bibi-haha /FineVideo-Prototype-Tokenized FineVideo-Prototype-Tokenized — Base Video Token Dataset Overview This dataset contains the base video tokenization output from the prototype pipeline, extracted from ~40K YouTube videos in the FineVideo dataset. Each video is tokenised into three modalities: Seed2 — 1 FPS semantic keyframe tokens (vocab: 8,192) Cosmos — every 8 frames spatial video tokens (vocab: 64,000) AVC-LM — every 8 frames H.264 BPE tokens (vocab: 8,192) This dataset does not contain 3D… See the full description on the dataset page: https://huggingface.co/datasets/bibi-haha/FineVideo-Prototype-Tokenized.textvideo-classification10K<n<100K0 likes6 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.