CoolFace
Datasetpublic

latkes/factprobe-replication-SUPERSEDED-stage2-counts-olmotok-v1

SUPERSEDED - do not use Renamed 2026-08-25. Use latkes/factprobe-replication-stage2-counts-canonical-v1 instead. It was counted with a name list missing each entity's canonical Wikidata name for 10,471 of 42,235 entities (24.8%). It is kept only so earlier numbers can be traced to where they came from. Nothing current should read it. factprobe-replication-stage2-counts-olmotok-v1 Occurrences of every probed entity name in the OLMo-2-7B mid-training corpus (the… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-SUPERSEDED-stage2-counts-olmotok-v1.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes19downloads
Dataset Card

SUPERSEDED - do not use

Renamed 2026-08-25. Use [`latkes/factprobe-replication-stage2-counts-canonical-v1`](https://huggingface.co/datasets/latkes/factprobe-replication-stage2-counts-canonical-v1) instead.

It was counted with a name list missing each entity's canonical Wikidata name for 10,471 of 42,235 entities (24.8%).

It is kept only so earlier numbers can be traced to where they came from. Nothing current should read it.


factprobe-replication-stage2-counts-olmotok-v1

Occurrences of every probed entity name in the OLMo-2-7B mid-training corpus (the 50B-token mix), counted as OLMo token sequences with a word boundary required at both ends of every match. Counted on Spark over all 576 published token files, 192 GB. Supersedes all earlier phase-2 counts: those were made with a tokenizer that cut text by GPT-2's rule instead of the rule AI2 tokenized the corpus with, which made 12,000 of the 164,952 names unfindable, including 'U.S.' See docs/tokenizer-splitting-rule.md.

Dataset Info

  • —Rows: 118725
  • —Columns: 6

Columns

ColumnTypeDescription
aliasValue('string')the entity name searched for
countleadspaceValue('int64')occurrences of the space-preceded form, i.e. the name in running text
count_bareValue('int64')occurrences of the bare form, i.e. after a line break, quote or bracket
countValue('int64')countleadspace + count_bare, the total occurrences of this name
corpusValue('string')which corpus was counted
constructValue('string')what exactly was counted: tokenizer, splitting rule, boundary rule

Generation Parameters

json
{
  "script_name": "count_npy_direct.py + merge_counts.py",
  "model": "OLMo-2-7B mid-training corpus (not a model run)",
  "description": "Occurrences of every probed entity name in the OLMo-2-7B mid-training corpus (the 50B-token mix), counted as OLMo token sequences with a word boundary required at both ends of every match. Counted on Spark over all 576 published token files, 192 GB. Supersedes all earlier phase-2 counts: those were made with a tokenizer that cut text by GPT-2's rule instead of the rule AI2 tokenized the corpus with, which made 12,000 of the 164,952 names unfindable, including 'U.S.' See docs/tokenizer-splitting-rule.md.",
  "experiment_name": "factprobe-replication",
  "cluster": "spark",
  "artifact_status": "final",
  "canary": false,
  "hyperparameters": {
    "tokenizer": "allenai/dolma2-tokenizer, pre-tokenizer taken from AI2's tokenizer.json",
    "matching": "ahocorasick_rs over raw token-id bytes, 4-byte aligned",
    "boundary_rule": "token after a match must start a new word; bare form must also begin one",
    "chunk_tokens": 25000000
  },
  "input_datasets": [
    "olmo-2-1124-7B stage-2 mix, 576 token files"
  ]
}

Usage

python
from datasets import load_dataset

dataset = load_dataset("latkes/factprobe-replication-stage2-counts-olmotok-v1", split="train")
print(f"Loaded {len(dataset)} rows")