factprobe-replication
factprobe-replication-negatives-allnames-v1
Plausible wrong answers, asked under every name (42,267,800 rows)
Status: final — 40 of 40 runs. Models present:
13b, 7b. Training stages present: s1, s2, s3, s4, s5.
What this fixes
A real fact is put to the model under the full cross product of the two
people's name lists, and counts as recognised if any one combination gets a
Yes. That is He et al.'s rule. Until 2026-08-26 the wrong answer it was
compared against was asked under one name per person, so the… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-negatives-allnames-v1.factprobe-replication-traj-b2-stage1-v1
factprobe-replication-traj-b2-stage1-v1
Checkpoint-trajectory probing: b2 few-shot recognition on 13 log-spaced stage-1 checkpoints per model, spouse+sibling, both templates, original alternating demos, fp16 pinned. Identity columns model_tag/revision/tokens_b injected from filenames. Complete checkpoint files only.
Dataset Info
Rows: 26315536
Columns: 18
Columns
Column
Type
Description
relation
Value('string')
No description provided… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-traj-b2-stage1-v1.factprobe-replication-asked-negatives-v1
Asking a model for a plausible wrong answer (100,776 questions)
Two prompts, put to gpt-5.1 for every subject surface form in the spouse
and sibling data, ten independent samples each:
(a) "Who is the spouse of X? Just the name, no explanation needed."
(b') "Name a person who could reasonably be mistaken for the spouse of X,
but is not the spouse of X. Just the name, no explanation needed."
50,388 surface forms across 10,592 entities, times two prompts, is
100,776 questions… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-asked-negatives-v1.factprobe-replication-stage1-fact-counts-v1
Correction, 2026-08-26
An earlier version of this card said the corpus states facts asymmetrically
in a way that tracks entity frequency, and gave 70.9% as the figure. That
number is right, and the generalisation drawn from it was wrong.
It holds for spouse and for no other relation. Recomputed across all four:
relation
pairs with a fact sentence
written more often with the MORE frequent entity first
spouse
1,265
70.9%
sibling
203
36.5% — the opposite
twinned town… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-stage1-fact-counts-v1.factprobe-replication-stagematched-7b-v1
Stage-matched probing of OLMo-2-7B (49,218,520 rows)
Five training stages (S1 end of pretraining, S2 released base, S3 SFT, S4 DPO,
S5 Instruct) x four Wikidata relations (P26 spouse, P3373 sibling, P190
twinnedTown, P47 bordersWith) x two phrasings (question, statement), each run
with 1:1 scrambled negatives drawn with a fixed seed so the negative set is
identical at every stage. Few-shot prompt held fixed across all stages; only
the model weights differ.
column
meaning… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-stagematched-7b-v1.factprobe-replication-stage1-counts-canonical-v1
factprobe-replication-stage1-counts-canonical-v1
Occurrences of every probed entity name in the OLMo-2 PRETRAINING corpus (olmo-mix-1124, 1,117 token files, 14.10 TiB), counted as OLMo token sequences with a word boundary required at each end, both written forms kept separate. SUPERSEDES factprobe-replication-stage1-counts-olmotok-v1, which was missing each entity's canonical name: the released triples name entities by their Wikidata aliases, a field that by construction… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-stage1-counts-canonical-v1.
