CoolFace
Datasetpublic

skandermoalla/qrpo-paper-mistral-sft-magpieair-armorm-temp1-ref50-offpolicy2random-armorm

qrpo-paper-mistral-sft-magpieair-armorm-temp1-ref50-offpolicy2random-armorm Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization). Part of the dataset collection for the paper Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions (https://arxiv.org/pdf/2507.08068).

sourceHugging Facemitupdated 10mo agoView on Hugging Face
0likes293downloads
3 commits on main
21cc46e10mo ago

Update README with code repo link

skandermoalla
5118b9710mo ago

Update dataset files: mistral-sft-magpieair-armorm-temp1-ref50-offpolicy2random-armorm

skandermoalla
0df586710mo ago

initial commit

skandermoalla