RichardErkhov/chujiezheng_-_Llama-3-Instruct-8B-SimPO-ExPO-gguf
Quantization made by Richard Erkhov.
Llama-3-Instruct-8B-SimPO-ExPO - GGUF
- Model creator: https://huggingface.co/chujiezheng/
- Original model: https://huggingface.co/chujiezheng/Llama-3-Instruct-8B-SimPO-ExPO/
Original model description: --- language:
- en license: llama3 ---
Llama-3-Instruct-8B-SimPO-ExPO
The extrapolated (ExPO) model based on `princeton-nlp/Llama-3-Instruct-8B-SimPO` and `meta-llama/Meta-Llama-3-8B-Instruct`, as in the "Weak-to-Strong Extrapolation Expedites Alignment" paper.
Specifically, we obtain this model by extrapolating (alpha = 0.3) from the weights of the SFT and DPO/RLHF checkpoints, achieving superior alignment with human preference.
This extrapolated model achieves the 40.6% win rate and 45.8% LC win rate on AlpacaEval 2.0, outperforming the original Llama-3-Instruct-8B-SimPO's 40.5% and 44.7%, respectively.
Evaluation Results
Evaluation results on the AlpacaEval 2.0 benchmark (you can find the evaluation outputs on the official GitHub repo):
Evaluation results on the MT-Bench benchmark (you can find the evaluation outputs on the official GitHub repo):
