CoolFace
Datasetpublic

LorMolf/SPSD-Variants-opsd

SPSD-Variants-opsd Grounded on-policy self-distillation (OPSD) teacher-context dataset over 45 board-game rule variants (5 families × 9: connect4, domineering, simplified_first_attack, simplified_othello, tic_tac_chess), derived from trained MuZero/EfficientZero checkpoints (plan-528 v2). Each row is a decision-state task (a move choice or one of six auxiliary state-QA tasks). The privileged_context is the teacher signal: grounded natural-language reasoning that discovers the… See the full description on the dataset page: https://huggingface.co/datasets/LorMolf/SPSD-Variants-opsd.

sourceHugging Faceotherupdated 23d agoView on Hugging Face
0likes59downloads
settings

This repository belongs to LorMolf on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameSPSD-Variants-opsd
visibilitypublic
licenceother
gatedno
ownerLorMolf
Account settings
LorMolf/SPSD-Variants-opsd · CoolFace