chujiezheng/Llama3-70B-Chinese-Chat-ExPO
Llama3-8B-Chinese-Chat-ExPO
The extrapolated (ExPO) model based on `shenzhi-wang/Llama3-70B-Chinese-Chat` and `meta-llama/Meta-Llama-3-70B-Instruct`, as in the "Weak-to-Strong Extrapolation Expedites Alignment" paper.
Specifically, we obtain this model by extrapolating (alpha = 0.3) from the weights of the SFT and DPO/RLHF checkpoints, achieving superior alignment with human preference.
Note: This is an experimental model, as I have not comprehensively evaluated its Chinese ability. Unexpected issues may occur when we apply extrapolation to the DPO/RLHF alignment training for new languages (e.g., Chinese).
Evaluation Results
Evaluation results on the AlpacaEval 2.0 benchmark (you can find the evaluation outputs on the official GitHub repo):
Evaluation results on the MT-Bench benchmark (you can find the evaluation outputs on the official GitHub repo):
