CoolFace
Datasetpublic

latkes/factprobe-replication-stage1-cooccurrence-v1

factprobe-replication-stage1-cooccurrence-v1 How many documents of the OLMo-2 pretraining corpus contain BOTH names of a probed pair. Computed over all 1,117 token files (15.50 TB, 3.875 trillion tokens) for the 2,172,383 name pairs the model was probed about; 313,576 of them share at least one document. Replaces an earlier version measured on the 192 GB mid-training mix alone, where 80-92% of pairs never co-occurred and the quantity behaved as a yes/no flag rather than a graded… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-stage1-cooccurrence-v1.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes14downloads
3 commits on main
6a916ad1mo ago

Upload README.md with huggingface_hub

juand-r
cda43121mo ago

Upload dataset

juand-r
15621e71mo ago

initial commit

juand-r