CoolFace
Datasetpublic

LorMolf/SPSD-Variants-opsd

SPSD-Variants-opsd Grounded on-policy self-distillation (OPSD) teacher-context dataset over 45 board-game rule variants (5 families × 9: connect4, domineering, simplified_first_attack, simplified_othello, tic_tac_chess), derived from trained MuZero/EfficientZero checkpoints (plan-528 v2). Each row is a decision-state task (a move choice or one of six auxiliary state-QA tasks). The privileged_context is the teacher signal: grounded natural-language reasoning that discovers the… See the full description on the dataset page: https://huggingface.co/datasets/LorMolf/SPSD-Variants-opsd.

sourceHugging Faceotherupdated 24d agoView on Hugging Face
0likes59downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
LorMolf/SPSD-Variants-opsd · CoolFace