datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
urls-tokenized
URLs (tokenized)
ks46/urls-sampled run through a byte-level
BPE built for URLs, stored as flat uint16 token streams that memory-map
directly into a training loop.
Shards
512
URLs
18,729,786,698
Tokens
664,731,047,208
Vocabulary
8,192
Token dtype
uint16, little-endian
There is no parquet here and the dataset viewer will not render it. These
are raw token bins; see Reading the data below.
Layout
tokenizer/ the exact vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls-tokenized.TRM-modified-datamix-tokenized
TRM modified datamix (tokenized)
Pre-tokenized reasoning/pretraining mixture for from-scratch TRM (Tiny Recursive Model)
training, built by running data_io — the HRM-Text data
pipeline — verbatim on sapientinc/HRM-Text-data-io-cleaned-20260515, with three
deliberate, documented deviations (below).
It is emitted in the V1 tokenized dataset format (a single concatenated token pool +
per-epoch document indices) and is ready to stream directly into training — no re-tokenization.… See the full description on the dataset page: https://huggingface.co/datasets/m-ric/TRM-modified-datamix-tokenized.movie-screenplays-tokenized-dataset
Screenplay Corpus — Tokenized (GPT-2)
Pre-tokenized screenplay corpus used to train and evaluate the models in the GPT-2 Screenplay Fine-Tuning Study. Derived from Movie-Script-Database by Aveek Saha. Provided as tokenized JSON splits ready for direct consumption by a GPT-2 Trainer pipeline — no preprocessing required.
Dataset Description
This dataset contains approximately 94 million tokens of professionally formatted screenplay text, pre-tokenized using the… See the full description on the dataset page: https://huggingface.co/datasets/raghavnimbalkar/movie-screenplays-tokenized-dataset.oellm-longctx-tokenized-superlong-512k-1m-2m-v1
OELLM Superlong Long-Context Tokenized 512K/1M/2M v1
This dataset is a superlong-context continuation-training add-on for extending beyond 256K toward 1M-2M context windows.
The design goal is not simply longer packed text. At 1M-2M, the model needs collection-level continuity and explicit pressure to use very old evidence. This artifact therefore mixes real long structured/natural sources with a small multilingual full-span recall component.
Source families:
RFC Editor… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/oellm-longctx-tokenized-superlong-512k-1m-2m-v1.FineVideo-Prototype-Tokenized
FineVideo-Prototype-Tokenized — Base Video Token Dataset
Overview
This dataset contains the base video tokenization output from the prototype pipeline, extracted from ~40K YouTube videos in the FineVideo dataset.
Each video is tokenised into three modalities:
Seed2 — 1 FPS semantic keyframe tokens (vocab: 8,192)
Cosmos — every 8 frames spatial video tokens (vocab: 64,000)
AVC-LM — every 8 frames H.264 BPE tokens (vocab: 8,192)
This dataset does not contain 3D… See the full description on the dataset page: https://huggingface.co/datasets/EmpathicRobotics/FineVideo-Prototype-Tokenized.c4-en-tokenized
C4 English Tokenized Samples
This dataset contains tokenized English samples from the C4 (Colossal Clean Crawled Corpus) dataset for natural language processing (NLP) tasks.
The first 125 000 entries from the en split of allenai/c4
were tokenized using spaCy's en_core_web_sm model. Tokens joined with spaces.
Features
text: Original text from C4
tokenized: The tokenized and space-joined text
num_tokens: Number of tokens after tokenization
num_punct_tokens: Number of… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/c4-en-tokenized.FineVideo-Prototype-Tokenized
FineVideo-Prototype-Tokenized — Base Video Token Dataset
Overview
This dataset contains the base video tokenization output from the prototype pipeline, extracted from ~40K YouTube videos in the FineVideo dataset.
Each video is tokenised into three modalities:
Seed2 — 1 FPS semantic keyframe tokens (vocab: 8,192)
Cosmos — every 8 frames spatial video tokens (vocab: 64,000)
AVC-LM — every 8 frames H.264 BPE tokens (vocab: 8,192)
This dataset does not contain 3D… See the full description on the dataset page: https://huggingface.co/datasets/bibi-haha/FineVideo-Prototype-Tokenized.
