CoolFace
Apppublic

sotayamashita/repro-understanding-the-performance-gap-in-preference-learning-a-dichotomy-of-rlhf-and-dpo

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes
29 commits on main
a3dc3a42mo ago

Upload pages/executive-summary/page.md with huggingface_hub

sotayamashita
dcb3d052mo ago

Upload pages/conclusion/page.md with huggingface_hub

sotayamashita
06562952mo ago

Upload pages/claim-4-sparse-reward-estimation-has-a-statistical-advantage/page.md with huggingface_hub

sotayamashita
21de7b52mo ago

Upload logbook.json with huggingface_hub

sotayamashita
5dc4dd02mo ago

Upload pages/index.md with huggingface_hub

sotayamashita
ac042c92mo ago

Upload pages/claim-3-rlhf-dominates-dpo-under-a-weak-policy-class/page.md with huggingface_hub

sotayamashita
352d3952mo ago

Upload pages/claim-2-online-dpo-can-strictly-outperform-offline-dpo/page.md with huggingface_hub

sotayamashita
2039bf42mo ago

Upload pages/claim-1-isomorphic-classes-make-offline-rlhf-and-dpo-equivalent/page.md with huggingface_hub

sotayamashita
26afa9f2mo ago

Upload logbook.json with huggingface_hub

sotayamashita
52185092mo ago

Upload README.md with huggingface_hub

sotayamashita
e5102242mo ago

Upload logbook.json with huggingface_hub

sotayamashita
4b9450d2mo ago

Upload logbook.json with huggingface_hub

sotayamashita
75fd77d2mo ago

Upload workspace.json with huggingface_hub

sotayamashita
2e6532a2mo ago

Upload pages/executive-summary/page.md with huggingface_hub

sotayamashita
be8169b2mo ago

Upload pages/conclusion/page.md with huggingface_hub

sotayamashita
77c66a72mo ago

Upload logbook.json with huggingface_hub

sotayamashita
c9c2cba2mo ago

Upload pages/executive-summary/page.md with huggingface_hub

sotayamashita
4b86e962mo ago

Upload pages/conclusion/page.md with huggingface_hub

sotayamashita
ea215302mo ago

Upload logbook.json with huggingface_hub

sotayamashita
9bf1a5e2mo ago

Upload pages/executive-summary/page.md with huggingface_hub

sotayamashita
d4c60482mo ago

Upload pages/conclusion/page.md with huggingface_hub

sotayamashita
b37f1242mo ago

Upload pages/claim-4-sparse-reward-estimation-has-a-statistical-advantage/page.md with huggingface_hub

sotayamashita
b75e29f2mo ago

Upload logbook.json with huggingface_hub

sotayamashita
3cd84462mo ago

Upload README.md with huggingface_hub

sotayamashita
1f325bd2mo ago

Upload logbook.json with huggingface_hub

sotayamashita
6870e6b2mo ago

Upload README.md with huggingface_hub

sotayamashita
6b476112mo ago

Upload folder using huggingface_hub

sotayamashita
90dbf0c2mo ago

Update logbook: Reproduction: Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO

sotayamashita
bd9e1222mo ago

initial commit

sotayamashita