placeholder
locus-commit-pool-v1
Locus Commit Pool v1
Native Git history, preserved as replayable software changes
Commit message · complete selected before-state · unified patches · native object IDs · provenance · experimental labels
Locus Commit Pool v1 is a large evidence pool for studying and training on how real software changes. Each document represents one surviving single-parent, multi-file Git commit. It keeps the commit message, the selected files as they existed before the… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/locus-commit-pool-v1.pretrain-ultra-fineweb-mixNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
73,328,911,347 (73.3B)
Trainable tokens
73,328,911,347 (73.3B)
Documents
92,923,076
Shards
578
UTF-8 bytes
364,562,557,837
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-ultra-fineweb-mix.locus-repo-code-pool-v2
Locus repository code pool
The Locus repository code pool contains curated repository-version documents
for code-language-model research. Each row combines the useful files from one
Git repository snapshot into a single deterministic text document while
retaining source provenance, file order, language and test signals, health
evidence, and stable content hashes.
This release is a candidate acquisition pool, not a ready-made training
split. Use a globally deduplicated manifest… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/locus-repo-code-pool-v2.exp-pool-repository-code-dolma2-tokenized
Locus EXP Repository Code - Dolma 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-repository-code-dolma2-tokenized.pool-academic
Placeholder Labs Academic Pool
Public organic academic pretraining pool with 41,112,360 documents and
456,510,356,112 Dolma-2 tokens.
Common Pile peS2o Filtered multidisciplinary research papers
Dolma 3 olmOCR Science PDFs (score >=0.30, academic evidence, clean extraction)
Common Pile LibreTexts Filtered open textbook sections
quality_bin is null; source scores are retained only as source-specific evidence
labels: document type and publication year
exact dedup plus verified… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pool-academic.Nemotron-SFT-Science-v2-Sharded
Nemotron-SFT-Science-v2-Sharded
Byte-preserving sequential 100 MB JSONL shards of selected files from nvidia/Nemotron-SFT-Science-v2. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms.
Included files: vendor.jsonl, so.jsonl, rqa.jsonl, syn_mcq.jsonl.
No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records and all… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Nemotron-SFT-Science-v2-Sharded.
