CoolFace
Modelpublic

dteome/dpo-mistral-7b-ultrafeedback-binarized-preferences-cleaned-v0.2

sourceHugging Facemitupdated 3y agoView on Hugging Face
0likes13downloads
Model Card

Mistral-7B-ultrafeedback-binarized-preferences-cleaned-v0.2

DPO training of Mistral-7B-Instruct-v0.2 using the fixed and cleaned UltraFeedback dataset argilla/ultrafeedback-binarized-preferences-cleaned

MT-Bench evalution

Pairwise comparison with gpt-4-turbo-preview

modelwinlosstiewin_rateloss_ratewin_rate_adjusted
mistral-7b-instruct-v0.2-ultrafeedback6120790.381250.125000.628125
mistral-7b-instruct-v0.22061790.125000.381250.371875