latkes/factprobe-replication-stage2-counts-canonical-v1
factprobe-replication-stage2-counts-canonical-v1 Occurrences of every probed entity name in the OLMo-2-7B mid-training corpus (576 token files), counted as OLMo token sequences with a word boundary required at each end, both written forms kept separate. SUPERSEDES factprobe-replication-stage2-counts-olmotok-v1, which was missing each entity's canonical name: the released triples name entities by their Wikidata aliases, and Wikidata keeps the canonical name in a separate field. 9… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-stage2-counts-canonical-v1.
This repository belongs to latkes on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
