mariiojastu/lmsys_chat_80_EE_multi_turn
Dataset Description This dataset consists of 80 prompts translated into Estonian from the LMSYS-Chat-1M dataset. The prompts are drawn from four categories: planning and scheduling (planeerimine), specific format writing (kirjutamine), mathematics (matemaatika), and logical reasoning (arutlemine). Thirty prompts in this dataset are previous translations from smugri-mt-bench and are included either verbatim or with minor modifications. Prompt Structure All prompts… See the full description on the dataset page: https://huggingface.co/datasets/mariiojastu/lmsys_chat_80_EE_multi_turn.
Dataset Description
This dataset consists of 80 prompts translated into Estonian from the LMSYS-Chat-1M dataset. The prompts are drawn from four categories: planning and scheduling (planeerimine), specific format writing (kirjutamine), mathematics (matemaatika), and logical reasoning (arutlemine).
Thirty prompts in this dataset are previous translations from smugri-mt-bench and are included either verbatim or with minor modifications.
Prompt Structure
All prompts are multi-turn. A second turn was written for each prompt to introduce additional constraints, provide new information, or challenge the model to review its response.
Intended Use
The dataset is designed for human evaluation of model-generated responses. Tasks were selected to align with a predefined evaluation rubric.
Evaluation Protocol
Evaluators compare responses from two anonymous models and select a preferred response. If two responses are of equal quality, a tie may be selected.
Responses are evaluated along the following dimensions:
- Naturalness Which model's response sounds more natural in Estonian?
- Grammatical correctness Which model's response is grammatically more correct in Estonian?
- Sequential instruction following Which model is better at handling sequential instructions?
- Usefulness (applicable to writing and planning tasks only) Which model's response is more useful?
- Correctness (applicable to mathematics and reasoning tasks only) Which model's response is mathematically or logically more correct?
- Overall preference Which model's response do you prefer overall?
