chujiezheng/Mistral7B-PairRM-SPPO-ExPO
Mistral7B-PairRM-SPPO-ExPO
The extrapolated (ExPO) model based on `UCLA-AGI/Mistral7B-PairRM-SPPO` and `mistralai/Mistral-7B-Instruct-v0.2`, as in the "Weak-to-Strong Extrapolation Expedites Alignment" paper.
Specifically, we obtain this model by extrapolating (alpha = 0.3) from the weights of the SFT and DPO/RLHF checkpoints, achieving superior alignment with human preference.
This extrapolated model achieves the 35.4% win rate and 31.8% LC win rate on AlpacaEval 2.0, outperforming the original Mistral7B-PairRM-SPPO's 32.2% and 30.5%, respectively.
Evaluation Results
Evaluation results on the AlpacaEval 2.0 benchmark (you can find the evaluation outputs on the official GitHub repo):
Evaluation results on the MT-Bench benchmark (you can find the evaluation outputs on the official GitHub repo):
Open LLM Leaderboard Evaluation Results
Detailed results can be found here
