datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
locus-commit-pool-v1
Locus Commit Pool v1
Native Git history, preserved as replayable software changes
Commit message · complete selected before-state · unified patches · native object IDs · provenance · experimental labels
Locus Commit Pool v1 is a large evidence pool for studying and training on how real software changes. Each document represents one surviving single-parent, multi-file Git commit. It keeps the commit message, the selected files as they existed before the… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/locus-commit-pool-v1.pretrain-ultra-fineweb-mixNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
73,328,911,347 (73.3B)
Trainable tokens
73,328,911,347 (73.3B)
Documents
92,923,076
Shards
578
UTF-8 bytes
364,562,557,837
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-ultra-fineweb-mix.locus-repo-code-pool-v2
Locus repository code pool
The Locus repository code pool contains curated repository-version documents
for code-language-model research. Each row combines the useful files from one
Git repository snapshot into a single deterministic text document while
retaining source provenance, file order, language and test signals, health
evidence, and stable content hashes.
This release is a candidate acquisition pool, not a ready-made training
split. Use a globally deduplicated manifest… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/locus-repo-code-pool-v2.exp-pool-repository-code-dolma2-tokenized
Locus EXP Repository Code - Dolma 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-repository-code-dolma2-tokenized.pool-academic
Placeholder Labs Academic Pool
Public organic academic pretraining pool with 41,112,360 documents and
456,510,356,112 Dolma-2 tokens.
Common Pile peS2o Filtered multidisciplinary research papers
Dolma 3 olmOCR Science PDFs (score >=0.30, academic evidence, clean extraction)
Common Pile LibreTexts Filtered open textbook sections
quality_bin is null; source scores are retained only as source-specific evidence
labels: document type and publication year
exact dedup plus verified… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pool-academic.Nemotron-SFT-Science-v2-Sharded
Nemotron-SFT-Science-v2-Sharded
Byte-preserving sequential 100 MB JSONL shards of selected files from nvidia/Nemotron-SFT-Science-v2. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms.
Included files: vendor.jsonl, so.jsonl, rqa.jsonl, syn_mcq.jsonl.
No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records and all… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Nemotron-SFT-Science-v2-Sharded.pretrain-web-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
8,689,580,607 (8.7B)
Trainable tokens
8,689,580,607 (8.7B)
Documents
281,846
Shards
89
UTF-8 bytes
37,540,769,483
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-web-mix-long-context.pretrain-ultra-fineweb-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
1,367,358,024 (1.4B)
Trainable tokens
1,367,358,024 (1.4B)
Documents
48,077
Shards
73
UTF-8 bytes
6,386,740,105
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-ultra-fineweb-mix-long-context.pretrain-commits-v2-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
12,571,681,749 (12.6B)
Trainable tokens
4,460,160,435 (4.5B)
Documents
992,475
Shards
327
UTF-8 bytes
49,288,867,997
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-commits-v2-mix-long-context.pretrain-commits-v2-mixNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
79,318,557,379 (79.3B)
Trainable tokens
28,772,968,648 (28.8B)
Documents
31,431,846
Shards
696
UTF-8 bytes
310,412,170,445
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-commits-v2-mix.GLM-5.1-Reasoning-Main-Sharded
GLM-5.1 reasoning main — sequential shards
Byte-preserving 100 MB JSONL shards of the main subset from Jackrong/GLM-5.1-Reasoning-1M-Cleaned, by Jackrong, derived upstream from Kassadin88/GLM-5.1-1000000x. All credit for the original data and cleaning belongs to those publishers.
Only main.jsonl is included. No filtering, shuffling, schema changes, tokenization or truncation. Original JSON fields and complete records are preserved. Shards retain upstream order; random shard… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/GLM-5.1-Reasoning-Main-Sharded.pretrain-repository-v2-mixNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
77,467,752,674 (77.5B)
Trainable tokens
77,467,752,674 (77.5B)
Documents
24,040,993
Shards
740
UTF-8 bytes
317,294,681,915
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-repository-v2-mix.exp-pool-commit-code-dolma2-tokenized
Locus EXP Commit Code - Dolma 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-commit-code-dolma2-tokenized.Kimi-K2.5-Reasoning-General-Sharded
Kimi-K2.5-Reasoning-General-Sharded
Byte-preserving sequential 100 MB JSONL shards of selected files from Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms.
Included files: General-Distillation.jsonl.
No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records and all original fields… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Kimi-K2.5-Reasoning-General-Sharded.Nemotron-SFT-Agentic-v2-Selected-Sharded
Nemotron-SFT-Agentic-v2-Selected-Sharded
Byte-preserving sequential 100 MB JSONL shards of selected files from nvidia/Nemotron-SFT-Agentic-v2. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms.
Included files: data/tool_calling.jsonl, data/interactive_agent.jsonl.
No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Nemotron-SFT-Agentic-v2-Selected-Sharded.exp-pool-academic-dolma2-tokenized
Locus EXP Academic - OLMo 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-academic-dolma2-tokenized.pretrain-academic-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
43,694,042,993 (43.7B)
Trainable tokens
43,694,042,993 (43.7B)
Documents
1,001,557
Shards
373
UTF-8 bytes
183,279,720,921
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-academic-mix-long-context.pretrain-web-mixNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
73,646,210,143 (73.6B)
Trainable tokens
73,646,210,143 (73.6B)
Documents
61,059,647
Shards
590
UTF-8 bytes
341,537,872,441
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-web-mix.exp-pool-commit-code-raw
Locus EXP Commit Code - shuffled raw proxy pool
Deterministically shuffled commit-message and unified-diff documents with complete source metadata.
MANIFEST.json pins source identity, sampling policy, token budgets, and per-file checksums.
The paired Dolma-2-tokenized repository preserves prompt masking for reproducible proxy training.
exp-pool-olmo-web-dolma2-tokenized
Locus EXP OLMo Web - OLMo 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-olmo-web-dolma2-tokenized.pretrain-nemotron-math-mixNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
22,927,812,461 (22.9B)
Trainable tokens
22,927,812,461 (22.9B)
Documents
21,377,358
Shards
180
UTF-8 bytes
77,994,866,327
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-nemotron-math-mix.pool-encyclopedic
locus-1 encyclopedic pool
The encyclopedic pool for locus-1 - one of seven cluster
datasets that are mixed into the pretraining corpus. Every pool shares one schema, so a
mixture is a query rather than a rebuild.
6,498,683 documents, 7,314,613,330 tokens under allenai/dolma2-tokenizer@5292e5d6c0f40b67cc765fe41bec991cf4345b5c
Built from: HuggingFaceFW/finewiki (en @ 8bd13e72e6a0)
Build settings digest: cbd87eae472c6b30
Fingerprint compatibility digest: 6ee2a61c2f91d6a1… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pool-encyclopedic.exp-pool-encyclopedic-dolma2-tokenized
Locus EXP Encyclopedic - OLMo 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-encyclopedic-dolma2-tokenized.pretrain-academic-mixNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
67,177,671,409 (67.2B)
Trainable tokens
67,177,671,409 (67.2B)
Documents
11,031,348
Shards
531
UTF-8 bytes
301,588,326,802
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-academic-mix.exp-pool-nemotron-math-dolma2-tokenized
Locus EXP Nemotron Math - OLMo 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-nemotron-math-dolma2-tokenized.pretrain-encyclopedic-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
587,625,128 (587.6M)
Trainable tokens
587,625,128 (587.6M)
Documents
23,631
Shards
9
UTF-8 bytes
1,978,753,989
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-encyclopedic-mix-long-context.pretrain-repository-v2-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
130,158,375,824 (130.2B)
Trainable tokens
130,158,375,824 (130.2B)
Documents
2,578,578
Shards
1,168
UTF-8 bytes
535,260,344,241
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-repository-v2-mix-long-context.pretrain-encyclopedic-mixNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
7,321,112,013 (7.3B)
Trainable tokens
7,321,112,013 (7.3B)
Documents
6,498,683
Shards
58
UTF-8 bytes
27,544,569,761
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-encyclopedic-mix.pretrain-nemotron-math-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
1,446,296,439 (1.4B)
Trainable tokens
1,446,296,439 (1.4B)
Documents
42,379
Shards
23
UTF-8 bytes
4,965,563,314
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-nemotron-math-mix-long-context.placeholder_tiebe
