CoolFace
Datasetpublic

latkes/factprobe-replication-stage1-cooccurrence-v1

factprobe-replication-stage1-cooccurrence-v1 How many documents of the OLMo-2 pretraining corpus contain BOTH names of a probed pair. Computed over all 1,117 token files (15.50 TB, 3.875 trillion tokens) for the 2,172,383 name pairs the model was probed about; 313,576 of them share at least one document. Replaces an earlier version measured on the 192 GB mid-training mix alone, where 80-92% of pairs never co-occurred and the quantity behaved as a yes/no flag rather than a graded… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-stage1-cooccurrence-v1.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes14downloads

latkes/factprobe-replication-stage1-cooccurrence-v1 · main · files are served by the source, never re-hosted here