kaleinaNyan/kolibri-qwen2.5-7b-060225-rlhf-1
This is an instruction following model (based on Qwen2.5-7B base) optimized for Russian language.
The model was trained in two phases: SFT (training data composition is similar to kolibri-mistral-0427) and RLHF.
Current RLHF pipeline leads to degradation on IFEval, but the overall 'vibe' of the model improves significantly. I am currently investigating the causes of this degradation and exploring methods to further enhance instruction-following capabilities.
The model uses ChatML template. Adding a system prompt will likely improve the model's performance on your tasks (experiment with it).
Instruction following evals
The model was tested using the following benchmarks:
Russian LLM Arena (proxy eval via JINA)
The table below approximates Russian LLM Arena scores using the JINA Judge model. Take it with a grain of salt.
