latkes/factprobe-replication-stage1-cooccurrence-v1
factprobe-replication-stage1-cooccurrence-v1 How many documents of the OLMo-2 pretraining corpus contain BOTH names of a probed pair. Computed over all 1,117 token files (15.50 TB, 3.875 trillion tokens) for the 2,172,383 name pairs the model was probed about; 313,576 of them share at least one document. Replaces an earlier version measured on the 192 GB mid-training mix alone, where 80-92% of pairs never co-occurred and the quantity behaved as a yes/no flag rather than a graded… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-stage1-cooccurrence-v1.
014
Upload README.md with huggingface_hub
Upload dataset
initial commit
