CoolFace
Datasetpublic

homerquan/boardgamebench-answer-grpo

BoardGameBench Answer GRPO Dataset This dataset contains 10,000 BoardGameBench prompt/reward examples generated for GRPO-style reinforcement learning on board-game move selection. The final nemotron-boardgame-answer-lora-b4-safe-final adapter used this reviewed GRPO corpus after SFT and DPO. For that final pilot run, training used the first 512 examples from grpo_train.jsonl; the full 10k reviewed set is published here for reproducibility and follow-up training.… See the full description on the dataset page: https://huggingface.co/datasets/homerquan/boardgamebench-answer-grpo.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes29downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
homerquan/boardgamebench-answer-grpo · CoolFace