CoolFace
Modelpublic

chujiezheng/internlm2-chat-1_8b-ExPO

sourceHugging Faceotherupdated 2y agoView on Hugging Face
1likes26downloads
Model Card

internlm2-chat-1_8b-ExPO

The extrapolated (ExPO) model based on `internlm2-chat-1_8b` and `internlm/internlm2-chat-1_8b-sft`, as in the "Weak-to-Strong Extrapolation Expedites Alignment" paper.

Specifically, we obtain this model by extrapolating (alpha = 0.5) from the weights of the SFT and DPO/RLHF checkpoints, achieving superior alignment with human preference.

Evaluation Results

Evaluation results on the AlpacaEval 2.0 benchmark (you can find the evaluation outputs on the official GitHub repo):

Win Rate (Ori)LC Win Rate (Ori)Win Rate (+ ExPO)LC Win Rate (+ ExPO)
HuggingFaceH4/zephyr-7b-alpha6.7%10.0%10.6%13.6%
HuggingFaceH4/zephyr-7b-beta10.2%13.2%11.1%14.0%
berkeley-nest/Starling-LM-7B-alpha15.0%18.3%18.2%19.5%
Nexusflow/Starling-LM-7B-beta26.6%25.8%29.6%26.4%
snorkelai/Snorkel-Mistral-PairRM24.7%24.0%28.8%26.4%
RLHFlow/LLaMA3-iterative-DPO-final29.2%36.0%32.7%37.8%
internlm/internlm2-chat-1.8b3.8%4.0%5.2%4.3%
internlm/internlm2-chat-7b20.5%18.3%28.1%22.7%
internlm/internlm2-chat-20b36.1%24.9%46.2%27.2%
allenai/tulu-2-dpo-7b8.5%10.2%11.5%11.7%
allenai/tulu-2-dpo-13b11.2%15.5%15.6%17.6%
allenai/tulu-2-dpo-70b15.4%21.2%23.0%25.7%

Evaluation results on the MT-Bench benchmark (you can find the evaluation outputs on the official GitHub repo):

Original+ ExPO
HuggingFaceH4/zephyr-7b-alpha6.856.87
HuggingFaceH4/zephyr-7b-beta7.027.06
berkeley-nest/Starling-LM-7B-alpha7.827.91
Nexusflow/Starling-LM-7B-beta8.108.18
snorkelai/Snorkel-Mistral-PairRM7.637.69
RLHFlow/LLaMA3-iterative-DPO-final8.088.45
internlm/internlm2-chat-1.8b5.175.26
internlm/internlm2-chat-7b7.727.80
internlm/internlm2-chat-20b8.138.26
allenai/tulu-2-dpo-7b6.356.38
allenai/tulu-2-dpo-13b7.007.26
allenai/tulu-2-dpo-70b7.798.03