pm-25/clembench-rlvr-dataset
Clembench RLVR Dataset – Full (Wins + Losses) combines SFT-Final (https://huggingface.co/datasets/clembench-playpen/SFT-Final-Dataset) and DPO_dialogue (https://huggingface.co/datasets/clembench-playpen/DPO_dialogue). Info: "id": concat of "game"+"episode" "query": "string in llama-format" "reward": rewards given (1=="chosen" from dpo-dialogue and "success" from sft-final, 0== "rejected" from dpo-dialogue) "origin": marks origin of sample "player": kept for dpo-dialogue… See the full description on the dataset page: https://huggingface.co/datasets/pm-25/clembench-rlvr-dataset.
Clembench RLVR Dataset – Full (Wins + Losses)
combines SFT-Final (https://huggingface.co/datasets/clembench-playpen/SFT-Final-Dataset) and DPO_dialogue (https://huggingface.co/datasets/clembench-playpen/DPO_dialogue).
Info:<br> "id": concat of "game"+"episode" <br> "query": "string in llama-format" <br> "reward": rewards given (1=="chosen" from dpo-dialogue and "success" from sft-final, 0== "rejected" from dpo-dialogue) <br> "origin": marks origin of sample <br> "player": kept for dpo-dialogue samples, to maybe split palyer1 and player 2 later if dataleakage occurrs <br> "format": temporary col used to kepp track if "query"-contant is in right format <br>
created at 2025-08-05T11:42:25+00:00 via Kaggle.
