CoolFace
Datasetpublic

latkes/factprobe-replication-stage1-counts-canonical-v1

factprobe-replication-stage1-counts-canonical-v1 Occurrences of every probed entity name in the OLMo-2 PRETRAINING corpus (olmo-mix-1124, 1,117 token files, 14.10 TiB), counted as OLMo token sequences with a word boundary required at each end, both written forms kept separate. SUPERSEDES factprobe-replication-stage1-counts-olmotok-v1, which was missing each entity's canonical name: the released triples name entities by their Wikidata aliases, a field that by construction… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-stage1-counts-canonical-v1.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes13downloads
3 commits on main
09086671mo ago

Upload README.md with huggingface_hub

juand-r
f162e311mo ago

Upload dataset

juand-r
3ddc9a41mo ago

initial commit

juand-r