CoolFace
Datasetpublic

LorMolf/SPSD-Variants-opsd

SPSD-Variants-opsd Grounded on-policy self-distillation (OPSD) teacher-context dataset over 45 board-game rule variants (5 families × 9: connect4, domineering, simplified_first_attack, simplified_othello, tic_tac_chess), derived from trained MuZero/EfficientZero checkpoints (plan-528 v2). Each row is a decision-state task (a move choice or one of six auxiliary state-QA tasks). The privileged_context is the teacher signal: grounded natural-language reasoning that discovers the… See the full description on the dataset page: https://huggingface.co/datasets/LorMolf/SPSD-Variants-opsd.

sourceHugging Faceotherupdated 24d agoView on Hugging Face
0likes59downloads
4 commits on main
f707dc624d ago

dataset card

LorMolf
4c2841924d ago

grounded OPSD v2 corpus (45 variants; 213,406 train + 47,681 held-out; gzip)

LorMolf
e73ecff24d ago

dataset card

LorMolf
137b90d24d ago

initial commit

LorMolf