latkes/factprobe-replication-SUPERSEDED-stage2-counts-olmotok-v1
SUPERSEDED - do not use Renamed 2026-08-25. Use latkes/factprobe-replication-stage2-counts-canonical-v1 instead. It was counted with a name list missing each entity's canonical Wikidata name for 10,471 of 42,235 entities (24.8%). It is kept only so earlier numbers can be traced to where they came from. Nothing current should read it. factprobe-replication-stage2-counts-olmotok-v1 Occurrences of every probed entity name in the OLMo-2-7B mid-training corpus (the… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-SUPERSEDED-stage2-counts-olmotok-v1.
SUPERSEDED - do not use
Renamed 2026-08-25. Use [`latkes/factprobe-replication-stage2-counts-canonical-v1`](https://huggingface.co/datasets/latkes/factprobe-replication-stage2-counts-canonical-v1) instead.
It was counted with a name list missing each entity's canonical Wikidata name for 10,471 of 42,235 entities (24.8%).
It is kept only so earlier numbers can be traced to where they came from. Nothing current should read it.
factprobe-replication-stage2-counts-olmotok-v1
Occurrences of every probed entity name in the OLMo-2-7B mid-training corpus (the 50B-token mix), counted as OLMo token sequences with a word boundary required at both ends of every match. Counted on Spark over all 576 published token files, 192 GB. Supersedes all earlier phase-2 counts: those were made with a tokenizer that cut text by GPT-2's rule instead of the rule AI2 tokenized the corpus with, which made 12,000 of the 164,952 names unfindable, including 'U.S.' See docs/tokenizer-splitting-rule.md.
Dataset Info
- Rows: 118725
- Columns: 6
Columns
Generation Parameters
{
"script_name": "count_npy_direct.py + merge_counts.py",
"model": "OLMo-2-7B mid-training corpus (not a model run)",
"description": "Occurrences of every probed entity name in the OLMo-2-7B mid-training corpus (the 50B-token mix), counted as OLMo token sequences with a word boundary required at both ends of every match. Counted on Spark over all 576 published token files, 192 GB. Supersedes all earlier phase-2 counts: those were made with a tokenizer that cut text by GPT-2's rule instead of the rule AI2 tokenized the corpus with, which made 12,000 of the 164,952 names unfindable, including 'U.S.' See docs/tokenizer-splitting-rule.md.",
"experiment_name": "factprobe-replication",
"cluster": "spark",
"artifact_status": "final",
"canary": false,
"hyperparameters": {
"tokenizer": "allenai/dolma2-tokenizer, pre-tokenizer taken from AI2's tokenizer.json",
"matching": "ahocorasick_rs over raw token-id bytes, 4-byte aligned",
"boundary_rule": "token after a match must start a new word; bare form must also begin one",
"chunk_tokens": 25000000
},
"input_datasets": [
"olmo-2-1124-7B stage-2 mix, 576 token files"
]
}Usage
from datasets import load_dataset
dataset = load_dataset("latkes/factprobe-replication-stage2-counts-olmotok-v1", split="train")
print(f"Loaded {len(dataset)} rows")