caiotheodoro/plumb
Plumb Gold tasks and the held-out eval. The study is on the collection. Config n What benchmark 1000 Held-out eval, seed 777. Never in train. train_handseeded 223 Mix matched to the eval, including PASS. train_ornith 58 Ornith-1.5 proposals that passed the oracle. train_blended 281 Both of the above. Leakprobe vs benchmark: exact signature overlap 0. curriculum train n sw-recall precision exact hand-seeded 223 0.318 [0.290, 0.347] 0.308 [0.279… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/plumb.
evals: flat published rows + CIs (reproduces the post tables)
dataset card: add evals config
Cut repeated lecture; cards are TL;DR
Rewrite card: thesis, domain, and standalone context
Upload PATH_A_RESULTS.json with huggingface_hub
Upload folder using huggingface_hub
initial commit
