factprobe
factprobe-replication-negatives-allnames-v1
Plausible wrong answers, asked under every name (42,267,800 rows)
Status: final — 40 of 40 runs. Models present:
13b, 7b. Training stages present: s1, s2, s3, s4, s5.
What this fixes
A real fact is put to the model under the full cross product of the two
people's name lists, and counts as recognised if any one combination gets a
Yes. That is He et al.'s rule. Until 2026-08-26 the wrong answer it was
compared against was asked under one name per person, so the… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-negatives-allnames-v1.factprobe-replication-traj-b2-stage1-v1
factprobe-replication-traj-b2-stage1-v1
Checkpoint-trajectory probing: b2 few-shot recognition on 13 log-spaced stage-1 checkpoints per model, spouse+sibling, both templates, original alternating demos, fp16 pinned. Identity columns model_tag/revision/tokens_b injected from filenames. Complete checkpoint files only.
Dataset Info
Rows: 26315536
Columns: 18
Columns
Column
Type
Description
relation
Value('string')
No description provided… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-traj-b2-stage1-v1.factprobe-replication-asked-negatives-v1
Asking a model for a plausible wrong answer (100,776 questions)
Two prompts, put to gpt-5.1 for every subject surface form in the spouse
and sibling data, ten independent samples each:
(a) "Who is the spouse of X? Just the name, no explanation needed."
(b') "Name a person who could reasonably be mistaken for the spouse of X,
but is not the spouse of X. Just the name, no explanation needed."
50,388 surface forms across 10,592 entities, times two prompts, is
100,776 questions… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-asked-negatives-v1.factprobe-replication-generation-spouse-7b-base-v1
factprobe-replication-generation-spouse-7b-base-v1
Free-form spouse generation: for every P26 subject NAME form, the base 7B model was asked 'Who is the {spouse} of ? Answer with just the name:' with 4 in-context demos, and sampled 5 times with nucleus sampling (top_p=0.95, temperature=1.0, max_tokens=64, stop at newline). 28,815 subject names. Companion to the P(Yes) probing datasets — this is what the model GENERATES, not a yes/no score.
Dataset Info
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-generation-spouse-7b-base-v1.factprobe-replication-exposure-knows-trajectory-v1
factprobe-replication-exposure-knows-trajectory-v1
Does EXACT per-checkpoint cumulative corpus exposure predict whether OLMo-2 KNOWS a fact? Two measures per (checkpoint, relation): PAIRED (P(Yes|true) > P(Yes|hard-negative), the 'beats' defs) and MARGINAL (P(Yes|true), r2_p_true + per-feature Spearman). Checkpoints: after-phase1, base(s1+s2), SFT, DPO, RLVR + 7 pre-merge s2-anneal ingredients; both models; both relations (P26 spouse, P3373 sibling); surface/name level; all 7… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-exposure-knows-trajectory-v1.factprobe-replication-generation-spouse-13b-base-v1
factprobe-replication-generation-spouse-13b-base-v1
Free-form spouse generation: for every P26 subject NAME form, the base model was asked 'Who is the {spouse} of ? Answer with just the name:' with 4 in-context demos, and sampled 5 times with nucleus sampling (top_p=0.95, temperature=1.0, max_tokens=64, stop at newline). 28,815 subject names. Companion to the P(Yes) probing datasets — this is what the model GENERATES, not a yes/no score.
Dataset Info
Rows: 30796… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-generation-spouse-13b-base-v1.
