latkes/factprobe-replication-SUPERSEDED-olmomix-counts-v1
SUPERSEDED - do not use Renamed 2026-08-25. Use latkes/factprobe-replication-stage1-counts-canonical-v1 instead. It was produced by querying the infini-gram service, which indexes the corpus with the Llama-2 tokenizer, rather than by counting the corpus as OLMo token sequences. It also predates the canonical-name repair. It is kept only so earlier numbers can be traced to where they came from. Nothing current should read it. factprobe-replication-olmomix-counts-v1… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-SUPERSEDED-olmomix-counts-v1.
SUPERSEDED - do not use
Renamed 2026-08-25. Use [`latkes/factprobe-replication-stage1-counts-canonical-v1`](https://huggingface.co/datasets/latkes/factprobe-replication-stage1-counts-canonical-v1) instead.
It was produced by querying the infini-gram service, which indexes the corpus with the Llama-2 tokenizer, rather than by counting the corpus as OLMo token sequences. It also predates the canonical-name repair.
It is kept only so earlier numbers can be traced to where they came from. Nothing current should read it.
factprobe-replication-olmomix-counts-v1
Exact olmo-mix-1124 occurrence counts for all 164,949 alias strings of the 42,235 probed FactProbe entities (infini-gram v4olmo-mix-1124llama, word-boundary token-sequence counts), plus per-entity sums over each entity's alias set (rows with alias = _entity::<QID>). 42,233 entities fully counted; 2 entities have an alias with only approximate counts available (count null, approxdropped true).
Dataset Info
- Rows: 207187
- Columns: 4
Columns
Generation Parameters
{
"script_name": "count_lane.py (4 lanes across hosts) + upload_counts.py",
"model": "n/a (corpus statistics)",
"description": "Exact olmo-mix-1124 occurrence counts for all 164,949 alias strings of the 42,235 probed FactProbe entities (infini-gram v4_olmo-mix-1124_llama, word-boundary token-sequence counts), plus per-entity sums over each entity's alias set (rows with alias = __entity__::<QID>). 42,233 entities fully counted; 2 entities have an alias with only approximate counts available (count null, approx_dropped true).",
"hyperparameters": {
"pacing_s": 2.6,
"lanes": 4
},
"input_datasets": [
"KRR-Oxford/FactProbe Zenodo triples (alias lists)"
],
"experiment_name": "factprobe-replication",
"job_id": "local lanes + mll:66091",
"cluster": "local+mll",
"artifact_status": "final",
"canary": false
}Usage
from datasets import load_dataset
dataset = load_dataset("latkes/factprobe-replication-olmomix-counts-v1", split="train")
print(f"Loaded {len(dataset)} rows")