latkes/factprobe-replication-stage2-fact-counts-v1
factprobe-replication-stage2-fact-counts-v1 How often the 7B mid-training corpus states each fact in so many words -- 'Netherlands borders Germany' -- rather than merely naming both entities in one document. 383,132 sentences were searched: four phrasings per relation, both directions, each entity written with its canonical name. 1,305 occur at all, 6,660 occurrences in total. This corpus is 0.19 TiB; the same measurement over the 14 TiB pretraining corpus is being counted now… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-stage2-fact-counts-v1.
This repository belongs to latkes on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
