datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb-tokenized
FineWeb Tokenized
> 4 trillion tokens of the pre-tokenized data the 🌐 web has to offer
What is it?
This is a pre-tokenized version of the HuggingFaceFW/fineweb dataset (currently in-progress, tokenization of the ~15 trillion tokens corpus is ongoing). The data is being pre-processed and tokenized using the AnisoleAI BPE tokenizer (52,022 vocabulary size) and packed into compact uint16 Parquet shards.
By distributing the pre-tokenized corpus, we eliminate… See the full description on the dataset page: https://huggingface.co/datasets/anisoleai/fineweb-tokenized.Stack_Tokenizedfineweb-tokenized-fake
What is it?
It's similar to anisolai/fineweb-tokenized but fake.
I don't understand why I did that :)
WARNING:
WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN SOMEONE.
WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN SOMEONE.
WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN SOMEONE.
WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN… See the full description on the dataset page: https://huggingface.co/datasets/mondk/fineweb-tokenized-fake.Scientific_Research_Tokenized
NexaSci Scientific Research Tokenized
This dataset repository now holds the active NexaSci scientific pretraining reservoir, the NexaMat controller fine-tuning pack, and archived legacy reservoir builds. The current production reservoir is the 10B-token Apache Arrow release under nexasci_reservoir_v3_10b_prod_rust/.
Current Status
The active large-scale training artifact is:
nexasci_reservoir_v3_10b_prod_rust/
It was produced from the NexaSci 10B data-engineering campaign… See the full description on the dataset page: https://huggingface.co/datasets/AethronPhantom/Scientific_Research_Tokenized.Qwen-Terminal-ToolBench-Processed-Tokenized
Qwen Terminal ToolBench Processed Datasets
Qwen-family processed/template-applied and selected tokenized terminal datasets.
Contents
qwen35_2b_full_terminal_toolcall_processed_v1: 1011776 rows, 36.26 GiB, features: text
qwen35_2b_terminal_toolcall_fullconv_from_samecount_v1: 327383 rows, 25.11 GiB, features: text
qwen35_27b_hf_fsdp_1024: 15335 rows, 0.17 GiB, features: input_ids, attention_mask, labels
qwen35_27b_processed_2bdata: 23472 rows, 2.77 GiB, features:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/Qwen-Terminal-ToolBench-Processed-Tokenized.exp-pool-repository-code-dolma2-tokenized
Locus EXP Repository Code - Dolma 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-repository-code-dolma2-tokenized.carbon-tokenized-corpus
Dataset Summary
AINovice2005/carbon-tokenized-corpus is the tokenized representation of sequence intervals processed in the Carbon enrichment pipeline.
Schema
The current dataset contains the following fields:
Field
Type
Description
record_id
string
Source/reference sequence identifier
start
int64
Start coordinate of the sequence interval
end
int64
End coordinate of the sequence interval
token_ids
list
Integer token IDs produced by the tokenizer… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-tokenized-corpus.clt_gpt2_tokenized_control
Fresh multilingual GPT-2 CLT control data
Sequential, unshuffled control sample for CLT null experiments. For each language,
complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision
0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded
(including the complete document that crossed the threshold), after which complete
documents were retained until at least 100,000,000 tokens were collected.
Data are… See the full description on the dataset page: https://huggingface.co/datasets/abir-hr196/clt_gpt2_tokenized_control.climbmix-tokenized-20480-diloco
ClimbMix, retokenized and shuffled for three-worker DiLoCo
This is a document-preserving, three-way split of NVIDIA's
Nemotron-ClimbMix,
retokenized with a 20,480-entry byte-level BPE tokenizer. Each document ends in
<|endoftext|>. The Arrow IPC streams use transparent Zstandard buffer
compression. A deterministic whole-shard holdout is shared by every worker for
validation and is excluded from training.
Training part
Documents
Tokens
Files
Compressed size
000
15,709… See the full description on the dataset page: https://huggingface.co/datasets/Sambarboi/climbmix-tokenized-20480-diloco.urls-tokenized
URLs (tokenized)
ks46/urls-sampled run through a byte-level
BPE built for URLs, stored as flat uint16 token streams that memory-map
directly into a training loop.
Shards
512
URLs
18,729,786,698
Tokens
664,731,047,208
Vocabulary
8,192
Token dtype
uint16, little-endian
There is no parquet here and the dataset viewer will not render it. These
are raw token bins; see Reading the data below.
Layout
tokenizer/ the exact vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls-tokenized.Pt-Corpus-Instruct-tokenized
Portuguese-Corpus Instruct (tokenized)
Dataset Summary
This repository has a tokenized version (using the TeenyTinyLlama tokenizer) of the Portuguese-Corpus Instruct dataset. All sequences are 2048 tokens long. All sequences are 2048 tokens long. This dataset was used in "TeenyTinyLlama: open-source tiny language models trained in Brazilian Portuguese".
For more information, see the original dataset card.
Languages
Portuguese.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/Pt-Corpus-Instruct-tokenized.dolma-v1_7-305B-tokenized-llama3-nanosetTokenized (Llama 3) verison of NousResearch/dolma-v1_7-305B as a Nanotron dataset split into 10 GB chunks.
To download:
huggingface-cli download --repo-type dataset --local-dir dolma-v1_7-305B-tokenized-llama3-nanoset --local-dir-use-symlinks False NousResearch/dolma-v1_7-305B-tokenized-llama3-nanoset
To recombine:
cat dolma-v1_7-305B-tokenized-llama3-nanoset/dolma-v1_7-305B-tokenized-llama3-nanoset.npy.* > dolma-v1_7-305B-tokenized-llama3-nanoset.npy
rm -rf… See the full description on the dataset page: https://huggingface.co/datasets/emozilla/dolma-v1_7-305B-tokenized-llama3-nanoset.exp-pool-commit-code-dolma2-tokenized
Locus EXP Commit Code - Dolma 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-commit-code-dolma2-tokenized.damr-zyda2-64k-tokenized
DAMR Zyda2 64K Tokenized
This repository contains a deterministic tokenized representation of a
1,749,978,785,641-token subset of
Zyda-2 for language-model
pretraining.
Format
Files under data/ contain contiguous little-endian uint16 token IDs.
Concatenate shards in numeric order to reproduce the original stream. The
shared 64K BPE tokenizer is stored under tokenizer/tokenizer.json. Shard
manifests provide byte offsets, token counts, and SHA-256 checksums.
The… See the full description on the dataset page: https://huggingface.co/datasets/zee-drytis/damr-zyda2-64k-tokenized.fineweb_10BT_tokenized
Dataset card for FineWeb-Edu 10B tokenized dataset
This dataset contains tokenized texts from FineWeb-Edu sample-10B HuggingFaceFW/fineweb-edu.
The data was tokenized using the OpenAI's tiktoken tokenizer, and structured for efficient streaming and distributed (DDP) training.
Structure
The dataset follows Hugging Face’s recommended structure for efficient streaming in multi-GPU environments.
It consists of two splits, where each split contains a number of shards… See the full description on the dataset page: https://huggingface.co/datasets/nikolina-p/fineweb_10BT_tokenized.llm0to1-pt-tokenized-en-edu-2025
LLM0to1 사전학습 토큰화본 — 영어 교육(2025 덤프)
10B 규모 한/영 이중언어 LLM LLM0to1-10b 를 바닥부터 학습할 때 실제로 투입된
영어 교육(2025 덤프) 코퍼스의 토큰화본. 총 28종 / 138.4B 토큰.
왜 원문 텍스트가 아니라 토큰화본인가
이 코퍼스들은 공개 데이터셋을 스트리밍으로 받아 곧바로 토큰화했고 중간 텍스트를 보관하지 않았다.
따라서 이 .ds 파일이 학습에 들어간 데이터의 유일한 사본이다.
단점만 있는 건 아니다. 토큰화본은 학습 입력 그 자체이므로,
재토큰화 과정에서 생길 수 있는 차이 없이 학습을 그대로 재현할 수 있다.
원본 출처
HuggingFaceFW/fineweb-edu 2025 덤프
영어 벤치마크 하락에 대응해 뒤늦게 편입한 '다양성 영어' 보강분이다. g00~`g15` 는 원본 shard 를 균등 분할한 것으로, 서로 다른 문서 집합이다.… See the full description on the dataset page: https://huggingface.co/datasets/izlley2/llm0to1-pt-tokenized-en-edu-2025.clt_gpt2_tokenized_control
Fresh multilingual GPT-2 CLT control data
Sequential, unshuffled control sample for CLT null experiments. For each language,
complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision
0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded
(including the complete document that crossed the threshold), after which complete
documents were retained until at least 100,000,000 tokens were collected.
Data are… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/clt_gpt2_tokenized_control.llm0to1-pt-tokenized-code
LLM0to1 사전학습 토큰화본 — 코드
10B 규모 한/영 이중언어 LLM LLM0to1-10b 를 바닥부터 학습할 때 실제로 투입된
코드 코퍼스의 토큰화본. 총 27종 / 59.1B 토큰.
왜 원문 텍스트가 아니라 토큰화본인가
이 코퍼스들은 공개 데이터셋을 스트리밍으로 받아 곧바로 토큰화했고 중간 텍스트를 보관하지 않았다.
따라서 이 .ds 파일이 학습에 들어간 데이터의 유일한 사본이다.
단점만 있는 건 아니다. 토큰화본은 학습 입력 그 자체이므로,
재토큰화 과정에서 생길 수 있는 차이 없이 학습을 그대로 재현할 수 있다.
원본 출처
bigcode/starcoderdata(언어별 서브셋) · bigcode/commitpackft · bigcode/jupyter-code-text-pairs · deepmind/code_contests
code_c 와 code_c2 처럼 2 가 붙은 것은 정제기… See the full description on the dataset page: https://huggingface.co/datasets/izlley2/llm0to1-pt-tokenized-code.exp-pool-academic-dolma2-tokenized
Locus EXP Academic - OLMo 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-academic-dolma2-tokenized.oellm-longctx-tokenized-streamed-all-v2
OpenEuroLLM long-context Megatron streamed tokenized artifacts
This dataset contains Megatron-LM indexed data (.bin/.idx) uploaded shard by shard from longctx stream-upload. It is intended as a transport format for continued pretraining on LUMI or another training machine; it is not raw text.
Source dataset: HuggingFaceFW/finepdfs-edu
Tooling: openeuro-longctx-datamix
Run namespace: runs/lc16k_full_20260507/
Tokenizer type: HuggingFaceTokenizer
Tokenizer: OpenEuroLLM 256k… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/oellm-longctx-tokenized-streamed-all-v2.tiny-stories-tokenized-bpeexp-pool-olmo-web-dolma2-tokenized
Locus EXP OLMo Web - OLMo 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-olmo-web-dolma2-tokenized.DepthBench-FineWeb-Edu-100BT-tokenized
DepthBench FineWeb-Edu 100BT Tokenized
This repository contains the tokenized FineWeb-Edu 100BT sample used by
DepthBench pretraining experiments.
Splits
train/: 139 shards, 99,585,913,529 tokens, and 97,045,608 documents.
eval/: 013_00008, containing 234,993,701 tokens and 225,078 documents.
All remaining source shards are assigned to training. Each source document is
terminated by an EOS token before documents are concatenated.
Format
Each shard… See the full description on the dataset page: https://huggingface.co/datasets/YangyiH/DepthBench-FineWeb-Edu-100BT-tokenized.Pt-Corpus-tokenized
Portuguese-Corpus (tokenized)
Dataset Summary
This repository has a tokenized version (using the TeenyTinyLlama tokenizer) of the Portuguese-Corpus dataset. All sequences are 2048 tokens long. This dataset was used in "TeenyTinyLlama: open-source tiny language models trained in Brazilian Portuguese".
For more information, see the original dataset card.
Languages
Portuguese.
Dataset Structure
Data Instances
The dataset consists… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/Pt-Corpus-tokenized.exp-pool-encyclopedic-dolma2-tokenized
Locus EXP Encyclopedic - OLMo 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-encyclopedic-dolma2-tokenized.exp-pool-nemotron-math-dolma2-tokenized
Locus EXP Nemotron Math - OLMo 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-nemotron-math-dolma2-tokenized.long_corpus-0209_tokenized
long_corpus-0209_tokenized
Exact-deduplicated, tokenized, and shuffled pretraining mix.
Processing
Tokenizer: /workspace/good_tokenizer (vocab_size=60800)
Special tokens: BOS=<bos> (id 2), EOS=<eos> (id 3)
bos_eos = True: every example is [BOS] + content + [EOS]
min_tokens = 15 (including BOS/EOS)
max_tokens = 3072: content truncated to 3070 tokens, then BOS+EOS appended (3070 + eos + bos)
Dedup: exact match on stripped UTF-8 text (blake2s-128)
Shuffle:… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/long_corpus-0209_tokenized.DeepSeek-R1-Distill-Llama-8B-lmsys-openthoughts-tokenizedThis dataset is used for training Sparse Autoencoders (SAEs) to identify reasoning features in Large Language Models (LLMs), as described in the paper I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders.
Code for the paper is available at: https://github.com/AIRI-Institute/SAE-Reasoning
The dataset consists of tokenized text data used for training the SAEs.
dataset_info:
features:
name: tokens
sequence: int64
splits:
name:… See the full description on the dataset page: https://huggingface.co/datasets/andreuka18/DeepSeek-R1-Distill-Llama-8B-lmsys-openthoughts-tokenized.ultimate-code-tokenized
ultimate-code
Dataset Description
ultimate-code is a derived dataset built by combining and processing data from
nvidia/OpenCodeInstruct and
nvidia/OpenCodeGeneticInstruct.
It can be used to fine-tune LLMs for coding tasks.
Tokenized variant available: a pre-tokenized version of this dataset is available at
CodeForCodersYT/ultimate-code-tokenized.
Use that version if you want ready-to-train tokenized sequences instead of raw text.
Source Datasets &… See the full description on the dataset page: https://huggingface.co/datasets/CodeForCodersYT/ultimate-code-tokenized.fineweb10B-tokenized-custom
LittleTzu FineWeb-Edu Tokenized (Custom 65k Balanced)
Tokenized shards of FineWeb-Edu (HuggingFaceFW/fineweb-edu, config: sample-10BT) for language model pretraining.
This dataset stores a derived, tokenized representation of the original FineWeb-Edu corpus. It has been tokenized using LittleTzu's custom 65K balanced tokenizer, optimized for multi-domain training (English, multilingual text, math, and code) while maintaining a compact vocabulary footprint that fits within a… See the full description on the dataset page: https://huggingface.co/datasets/Neetree/fineweb10B-tokenized-custom.
