CoolFace
Datasetpublic

latkes/factprobe-replication-stage2-counts-canonical-v1

factprobe-replication-stage2-counts-canonical-v1 Occurrences of every probed entity name in the OLMo-2-7B mid-training corpus (576 token files), counted as OLMo token sequences with a word boundary required at each end, both written forms kept separate. SUPERSEDES factprobe-replication-stage2-counts-olmotok-v1, which was missing each entity's canonical name: the released triples name entities by their Wikidata aliases, and Wikidata keeps the canonical name in a separate field. 9… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-stage2-counts-canonical-v1.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes8downloads
settings

This repository belongs to latkes on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namefactprobe-replication-stage2-counts-canonical-v1
visibilitypublic
licencemit
gatedno
ownerlatkes
Account settings
latkes/factprobe-replication-stage2-counts-canonical-v1 · CoolFace