skandermoalla/qrpo-paper-mistral-nosft-ultrafeedback-armorm-temp1-ref50-offline-armorm
qrpo-paper-mistral-nosft-ultrafeedback-armorm-temp1-ref50-offline-armorm Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization). Part of the dataset collection for the paper Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions (https://arxiv.org/pdf/2507.08068).
060
This repository belongs to skandermoalla on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
qrpo-paper-mistral-nosft-ultrafeedback-armorm-temp1-ref50-offline-armorm
public
mit
no
skandermoalla
