datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tokenized_datasettokenized_C4temp-decoder-train-tokenizedtokenized-dataset-combineurls-tokenized
URLs (tokenized)
ks46/urls-sampled run through a byte-level
BPE built for URLs, stored as flat uint16 token streams that memory-map
directly into a training loop.
Shards
512
URLs
18,729,786,698
Tokens
664,731,047,208
Vocabulary
8,192
Token dtype
uint16, little-endian
There is no parquet here and the dataset viewer will not render it. These
are raw token bins; see Reading the data below.
Layout
tokenizer/ the exact vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls-tokenized.jora_corpus1_FR_tokenized_128ksokoban-10k-vjepa2-tokenizedLWT-2.5B-model-6-tokenizedholistic-tokenizedLWT-2.5B-model-5-tokenizedLWT-2.5B-model-7-tokenizedLWT-2.5B-model-4-tokenizedmistral_tokenized_2048_fixed_shardsLWT-2.5B-model-8-tokenizedLWT-2.5B-model-1-tokenizedLWT-2.5B-model-3-tokenizedLWT-2.5B-model-2-tokenizedLWT-2.5B-stage0-tokenizedTRM-modified-datamix-tokenized
TRM modified datamix (tokenized)
Pre-tokenized reasoning/pretraining mixture for from-scratch TRM (Tiny Recursive Model)
training, built by running data_io — the HRM-Text data
pipeline — verbatim on sapientinc/HRM-Text-data-io-cleaned-20260515, with three
deliberate, documented deviations (below).
It is emitted in the V1 tokenized dataset format (a single concatenated token pool +
per-epoch document indices) and is ready to stream directly into training — no re-tokenization.… See the full description on the dataset page: https://huggingface.co/datasets/m-ric/TRM-modified-datamix-tokenized.miscellaneous_yt_chunked_tokenized
miscellaneous_yt_chunked_48k_tokenized
This is a gated Uzbek tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_tokenized.swe_multiturn-claude-trace-tokenized-qwen3.5
SWE-bench Multi-turn Trace for LLMServingSim2
Derived from
LMCache/Agentic-Traces,
filtered to the claude-sonnet-4-6 + swebench subset, then re-tokenized
with Qwen/Qwen3.5-122B-A10B
to drive an LLM serving simulator with realistic per-turn prefix-cache reuse
and strict turn-by-turn dependency enforcement.
110 sessions, 2,559 turns
Average input 20.2K tokens per turn (max 65.9K)
Average output 288.7 tokens per turn
Per-session prefix-share ratio: 0.939 (turn N input is ~94% the same… See the full description on the dataset page: https://huggingface.co/datasets/noddu/swe_multiturn-claude-trace-tokenized-qwen3.5.code_search_net_python_tokenizedmovie-screenplays-tokenized-dataset
Screenplay Corpus — Tokenized (GPT-2)
Pre-tokenized screenplay corpus used to train and evaluate the models in the GPT-2 Screenplay Fine-Tuning Study. Derived from Movie-Script-Database by Aveek Saha. Provided as tokenized JSON splits ready for direct consumption by a GPT-2 Trainer pipeline — no preprocessing required.
Dataset Description
This dataset contains approximately 94 million tokens of professionally formatted screenplay text, pre-tokenized using the… See the full description on the dataset page: https://huggingface.co/datasets/raghavnimbalkar/movie-screenplays-tokenized-dataset.fineweb-edu-100b-smollmv2-tokenizedoellm-longctx-tokenized-superlong-512k-1m-2m-v1
OELLM Superlong Long-Context Tokenized 512K/1M/2M v1
This dataset is a superlong-context continuation-training add-on for extending beyond 256K toward 1M-2M context windows.
The design goal is not simply longer packed text. At 1M-2M, the model needs collection-level continuity and explicit pressure to use very old evidence. This artifact therefore mixes real long structured/natural sources with a small multilingual full-span recall component.
Source families:
RFC Editor… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/oellm-longctx-tokenized-superlong-512k-1m-2m-v1.tokenized-crosscoder-qwen3-hhrf-dpo-pt
Tokenized CrossCoder Dataset (Qwen3 + HH-RLHF + DPO)
Description
This dataset contains tokenized preference pairs from Anthropic's HH-RLHF and LMSys Chatbot Arena conversations, preprocessed for Direct Preference Optimization (DPO) training using the Qwen3-8B-Base tokenizer.
Dataset Details
Tokenizer: Qwen/Qwen3-8B-Base
Max Sequence Length: 1024 tokens
Format: PyTorch tensors (.pt file)
Task: Direct Preference Optimization (DPO)
Sources:
Anthropic/hh-rlhf… See the full description on the dataset page: https://huggingface.co/datasets/Kasmic/tokenized-crosscoder-qwen3-hhrf-dpo-pt.PEFT_tokenized_datasetstokenizedfileshowdown-dojo-tokenizedaudiobook_chunked_tokenized
audiobook_chunked_tokenized
This is a gated Uzbek tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized speech training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_tokenized.
