CoolFace
Modelpublic

kaleinaNyan/kolibri-qwen2.5-7b-060225-rlhf-1

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes5downloads
Model Card

This is an instruction following model (based on Qwen2.5-7B base) optimized for Russian language.

The model was trained in two phases: SFT (training data composition is similar to kolibri-mistral-0427) and RLHF.

Current RLHF pipeline leads to degradation on IFEval, but the overall 'vibe' of the model improves significantly. I am currently investigating the causes of this degradation and exploring methods to further enhance instruction-following capabilities.

The model uses ChatML template. Adding a system prompt will likely improve the model's performance on your tasks (experiment with it).

Instruction following evals

The model was tested using the following benchmarks:

Eval nameStrict ValueLoose Value
Avg.43.0049.17
ifeval-prompt-level38.6346.21
ifeval-instruction-level51.2057.5
ru-ifeval-prompt-level35.3040.48
ru-ifeval-instruction-level46.8852.52

Russian LLM Arena (proxy eval via JINA)

The table below approximates Russian LLM Arena scores using the JINA Judge model. Take it with a grain of salt.

Model NameScore95% CIAvg Tokens
gpt-4-1106-preview82.8(-2.8, 2.6)541
gpt-4o-mini75.3(-2.2, 2.8)448
qwen-2.5-72b-it73.1(-3.0, 3.1)557
gemma-2-9b-it-sppo-iter370.6(-3.7, 3.0)509
gemma-2-27b-it68.7(-2.9, 3.8)472
t-lite-instruct-0.167.5(-4.2, 2.7)810
gemma-2-9b-it67.0(-3.0, 3.8)459
suzume-llama-3-8B-multilingual-orpo-borda-half62.4(-3.0, 3.3)682
glm-4-9b-chat61.5(-3.9, 3.3)568
phi-3-medium-4k-instruct60.4(-3.8, 3.6)566
sfr-iterative-dpo-llama-3-8b-r57.2(-3.8, 4.0)516
kolibri-qwen2.5-7b-060225-rlhf-155.4(-3.1, 4.4)383
c4ai-command-r-v0155.0(-3.7, 4.4)529
suzume-llama-3-8b-multilingual51.9(-3.1, 3.4)641
mistral-nemo-instruct-240751.9(-3.0, 3.0)403
yandexgptpro50.3(-3.5, 3.0)345
gpt-3.5-turbo-012550.0(0.0, 0.0)220
hermes-2-theta-llama-3-8b49.3(-3.2, 3.7)485
starling-lm-7b-beta48.3(-3.7, 3.9)629
llama-3-8b-saiga-suzume-ties47.9(-3.9, 5.0)763
llama-3-smaug-8b47.6(-4.3, 2.9)524
vikhr-it-5.4-fp16-orpo-v246.8(-2.4, 2.2)379
aya-23-8b46.1(-3.3, 3.6)554
saiga_llama3_8b_v644.8(-2.9, 3.2)471
qwen2-7b-instruct43.6(-3.5, 3.0)340
vikhr-it-5.2-fp16-cp43.6(-3.6, 3.3)543
openchat-3.5-010642.8(-2.5, 3.8)492
kolibri-mistral-0427-upd42.3(-4.1, 4.0)551
paralex-llama-3-8b-sft41.8(-3.7, 3.9)688
llama-3-instruct-8b-sppo-iter341.7(-4.0, 3.6)502
gpt-3.5-turbo-110641.5(-2.7, 2.5)191
mistral-7b-instruct-v0.341.1(-4.1, 2.9)469
gigachat_pro40.9(-3.2, 2.8)294
openchat-3.6-8b-2024052239.1(-2.9, 3.8)428
vikhr-it-5.3-fp16-32k38.8(-3.2, 3.3)519
hermes-2-pro-llama-3-8b38.4(-3.9, 3.9)463
kolibri-vikhr-mistral-042734.5(-2.9, 3.1)489
vikhr-it-5.3-fp1633.5(-3.0, 3.8)523
llama-3-instruct-8b-simpo32.7(-3.2, 2.7)417
meta-llama-3-8b-instruct32.1(-3.6, 4.2)450
neural-chat-7b-v3-325.9(-3.1, 3.2)927
gigachat_lite25.4(-3.5, 2.7)276
snorkel-mistral-pairrm-dpo10.3(-2.3, 2.6)773
storm-7b3.7(-1.9, 1.7)419