CNCL-Penn-State/CrPO-RM-gemma-2-9b-it
010
CrPO-RM-gemma-2-9b-it
This is the reward model (gemma-2-9b-it model finetuned on the MuCE dataset) used to measure the quality of model generations in the Creative Preference Optimization paper.
Citation
@misc{ismayilzada2025creativepreferenceoptimization,
title={Creative Preference Optimization},
author={Mete Ismayilzada and Antonio Laverghetta Jr. and Simone A. Luchini and Reet Patel and Antoine Bosselut and Lonneke van der Plas and Roger E. Beaty},
year={2025},
eprint={2505.14442},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2505.14442},
}