datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pretrain-ultra-fineweb-mixNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
73,328,911,347 (73.3B)
Trainable tokens
73,328,911,347 (73.3B)
Documents
92,923,076
Shards
578
UTF-8 bytes
364,562,557,837
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-ultra-fineweb-mix.exp-pool-repository-code-dolma2-tokenized
Locus EXP Repository Code - Dolma 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-repository-code-dolma2-tokenized.Nemotron-SFT-Science-v2-Sharded
Nemotron-SFT-Science-v2-Sharded
Byte-preserving sequential 100 MB JSONL shards of selected files from nvidia/Nemotron-SFT-Science-v2. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms.
Included files: vendor.jsonl, so.jsonl, rqa.jsonl, syn_mcq.jsonl.
No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records and all… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Nemotron-SFT-Science-v2-Sharded.pretrain-web-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
8,689,580,607 (8.7B)
Trainable tokens
8,689,580,607 (8.7B)
Documents
281,846
Shards
89
UTF-8 bytes
37,540,769,483
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-web-mix-long-context.pretrain-ultra-fineweb-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
1,367,358,024 (1.4B)
Trainable tokens
1,367,358,024 (1.4B)
Documents
48,077
Shards
73
UTF-8 bytes
6,386,740,105
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-ultra-fineweb-mix-long-context.pretrain-commits-v2-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
12,571,681,749 (12.6B)
Trainable tokens
4,460,160,435 (4.5B)
Documents
992,475
Shards
327
UTF-8 bytes
49,288,867,997
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-commits-v2-mix-long-context.pretrain-commits-v2-mixNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
79,318,557,379 (79.3B)
Trainable tokens
28,772,968,648 (28.8B)
Documents
31,431,846
Shards
696
UTF-8 bytes
310,412,170,445
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-commits-v2-mix.GLM-5.1-Reasoning-Main-Sharded
GLM-5.1 reasoning main — sequential shards
Byte-preserving 100 MB JSONL shards of the main subset from Jackrong/GLM-5.1-Reasoning-1M-Cleaned, by Jackrong, derived upstream from Kassadin88/GLM-5.1-1000000x. All credit for the original data and cleaning belongs to those publishers.
Only main.jsonl is included. No filtering, shuffling, schema changes, tokenization or truncation. Original JSON fields and complete records are preserved. Shards retain upstream order; random shard… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/GLM-5.1-Reasoning-Main-Sharded.pretrain-repository-v2-mixNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
77,467,752,674 (77.5B)
Trainable tokens
77,467,752,674 (77.5B)
Documents
24,040,993
Shards
740
UTF-8 bytes
317,294,681,915
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-repository-v2-mix.exp-pool-commit-code-dolma2-tokenized
Locus EXP Commit Code - Dolma 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-commit-code-dolma2-tokenized.Kimi-K2.5-Reasoning-General-Sharded
Kimi-K2.5-Reasoning-General-Sharded
Byte-preserving sequential 100 MB JSONL shards of selected files from Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms.
Included files: General-Distillation.jsonl.
No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records and all original fields… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Kimi-K2.5-Reasoning-General-Sharded.Nemotron-SFT-Agentic-v2-Selected-Sharded
Nemotron-SFT-Agentic-v2-Selected-Sharded
Byte-preserving sequential 100 MB JSONL shards of selected files from nvidia/Nemotron-SFT-Agentic-v2. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms.
Included files: data/tool_calling.jsonl, data/interactive_agent.jsonl.
No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Nemotron-SFT-Agentic-v2-Selected-Sharded.exp-pool-academic-dolma2-tokenized
Locus EXP Academic - OLMo 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-academic-dolma2-tokenized.pretrain-academic-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
43,694,042,993 (43.7B)
Trainable tokens
43,694,042,993 (43.7B)
Documents
1,001,557
Shards
373
UTF-8 bytes
183,279,720,921
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-academic-mix-long-context.pretrain-web-mixNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
73,646,210,143 (73.6B)
Trainable tokens
73,646,210,143 (73.6B)
Documents
61,059,647
Shards
590
UTF-8 bytes
341,537,872,441
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-web-mix.exp-pool-commit-code-raw
Locus EXP Commit Code - shuffled raw proxy pool
Deterministically shuffled commit-message and unified-diff documents with complete source metadata.
MANIFEST.json pins source identity, sampling policy, token budgets, and per-file checksums.
The paired Dolma-2-tokenized repository preserves prompt masking for reproducible proxy training.
exp-pool-olmo-web-dolma2-tokenized
Locus EXP OLMo Web - OLMo 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-olmo-web-dolma2-tokenized.pretrain-nemotron-math-mixNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
22,927,812,461 (22.9B)
Trainable tokens
22,927,812,461 (22.9B)
Documents
21,377,358
Shards
180
UTF-8 bytes
77,994,866,327
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-nemotron-math-mix.pool-encyclopedic
locus-1 encyclopedic pool
The encyclopedic pool for locus-1 - one of seven cluster
datasets that are mixed into the pretraining corpus. Every pool shares one schema, so a
mixture is a query rather than a rebuild.
6,498,683 documents, 7,314,613,330 tokens under allenai/dolma2-tokenizer@5292e5d6c0f40b67cc765fe41bec991cf4345b5c
Built from: HuggingFaceFW/finewiki (en @ 8bd13e72e6a0)
Build settings digest: cbd87eae472c6b30
Fingerprint compatibility digest: 6ee2a61c2f91d6a1… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pool-encyclopedic.exp-pool-encyclopedic-dolma2-tokenized
Locus EXP Encyclopedic - OLMo 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-encyclopedic-dolma2-tokenized.exp-pool-nemotron-math-dolma2-tokenized
Locus EXP Nemotron Math - OLMo 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-nemotron-math-dolma2-tokenized.pretrain-encyclopedic-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
587,625,128 (587.6M)
Trainable tokens
587,625,128 (587.6M)
Documents
23,631
Shards
9
UTF-8 bytes
1,978,753,989
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-encyclopedic-mix-long-context.pretrain-repository-v2-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
130,158,375,824 (130.2B)
Trainable tokens
130,158,375,824 (130.2B)
Documents
2,578,578
Shards
1,168
UTF-8 bytes
535,260,344,241
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-repository-v2-mix-long-context.pretrain-encyclopedic-mixNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
7,321,112,013 (7.3B)
Trainable tokens
7,321,112,013 (7.3B)
Documents
6,498,683
Shards
58
UTF-8 bytes
27,544,569,761
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-encyclopedic-mix.pretrain-nemotron-math-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
1,446,296,439 (1.4B)
Trainable tokens
1,446,296,439 (1.4B)
Documents
42,379
Shards
23
UTF-8 bytes
4,965,563,314
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-nemotron-math-mix-long-context.placeholder_tiebeexp-pool-olmo-web-raw
olmo_web shuffled raw proxy pool
Documents remain raw UTF-8 with complete source metadata.
Dolma-2 tokenization and 16K packing are intentionally deferred to training.
MANIFEST.json pins source identity, shuffle policy, capacity, and checksums.
exp-pool-academic-raw
academic shuffled raw proxy pool
Documents remain raw UTF-8 with complete source metadata.
Dolma-2 tokenization and 16K packing are intentionally deferred to training.
MANIFEST.json pins source identity, shuffle policy, capacity, and checksums.
placeholder_tiebe2exp-pool-repository-code-raw
Locus EXP Repository Code - shuffled raw proxy pool
Deterministically shuffled repository-code documents with complete source metadata.
MANIFEST.json pins source identity, sampling policy, token budgets, and per-file checksums.
The paired Dolma-2-tokenized repository is ready for reproducible proxy training.
