CoolFace
Apppublic

sotayamashita/repro-understanding-the-performance-gap-in-preference-learning-a-dichotomy-of-rlhf-and-dpo

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes
App README

Reproduction: Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO

An open experiment logbook, published with Trackio.