CoolFace
Datasetpublic

latkes/factprobe-replication-SUPERSEDED-olmomix-counts-v1

SUPERSEDED - do not use Renamed 2026-08-25. Use latkes/factprobe-replication-stage1-counts-canonical-v1 instead. It was produced by querying the infini-gram service, which indexes the corpus with the Llama-2 tokenizer, rather than by counting the corpus as OLMo token sequences. It also predates the canonical-name repair. It is kept only so earlier numbers can be traced to where they came from. Nothing current should read it. factprobe-replication-olmomix-counts-v1… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-SUPERSEDED-olmomix-counts-v1.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes24downloads
Dataset Card

SUPERSEDED - do not use

Renamed 2026-08-25. Use [`latkes/factprobe-replication-stage1-counts-canonical-v1`](https://huggingface.co/datasets/latkes/factprobe-replication-stage1-counts-canonical-v1) instead.

It was produced by querying the infini-gram service, which indexes the corpus with the Llama-2 tokenizer, rather than by counting the corpus as OLMo token sequences. It also predates the canonical-name repair.

It is kept only so earlier numbers can be traced to where they came from. Nothing current should read it.


factprobe-replication-olmomix-counts-v1

Exact olmo-mix-1124 occurrence counts for all 164,949 alias strings of the 42,235 probed FactProbe entities (infini-gram v4olmo-mix-1124llama, word-boundary token-sequence counts), plus per-entity sums over each entity's alias set (rows with alias = _entity::<QID>). 42,233 entities fully counted; 2 entities have an alias with only approximate counts available (count null, approxdropped true).

Dataset Info

  • —Rows: 207187
  • —Columns: 4

Columns

ColumnTypeDescription
aliasValue('string')Alias string as counted, or _entity_::<QID> for entity-level sums
countValue('int64')Exact occurrence count in olmo-mix-1124 (null if only approximate)
approx_droppedValue('bool')True if the API returned an approximate count (dropped) or the entity sum is incomplete
indexValue('string')infini-gram index (or entitysumoveraliasset for sum rows)

Generation Parameters

json
{
  "script_name": "count_lane.py (4 lanes across hosts) + upload_counts.py",
  "model": "n/a (corpus statistics)",
  "description": "Exact olmo-mix-1124 occurrence counts for all 164,949 alias strings of the 42,235 probed FactProbe entities (infini-gram v4_olmo-mix-1124_llama, word-boundary token-sequence counts), plus per-entity sums over each entity's alias set (rows with alias = __entity__::<QID>). 42,233 entities fully counted; 2 entities have an alias with only approximate counts available (count null, approx_dropped true).",
  "hyperparameters": {
    "pacing_s": 2.6,
    "lanes": 4
  },
  "input_datasets": [
    "KRR-Oxford/FactProbe Zenodo triples (alias lists)"
  ],
  "experiment_name": "factprobe-replication",
  "job_id": "local lanes + mll:66091",
  "cluster": "local+mll",
  "artifact_status": "final",
  "canary": false
}

Usage

python
from datasets import load_dataset

dataset = load_dataset("latkes/factprobe-replication-olmomix-counts-v1", split="train")
print(f"Loaded {len(dataset)} rows")