CoolFace
Datasetpublic

mariiojastu/lmsys_chat_80_EE_multi_turn

Dataset Description This dataset consists of 80 prompts translated into Estonian from the LMSYS-Chat-1M dataset. The prompts are drawn from four categories: planning and scheduling (planeerimine), specific format writing (kirjutamine), mathematics (matemaatika), and logical reasoning (arutlemine). Thirty prompts in this dataset are previous translations from smugri-mt-bench and are included either verbatim or with minor modifications. Prompt Structure All prompts… See the full description on the dataset page: https://huggingface.co/datasets/mariiojastu/lmsys_chat_80_EE_multi_turn.

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes19downloads
Dataset Card

Dataset Description

This dataset consists of 80 prompts translated into Estonian from the LMSYS-Chat-1M dataset. The prompts are drawn from four categories: planning and scheduling (planeerimine), specific format writing (kirjutamine), mathematics (matemaatika), and logical reasoning (arutlemine).

Thirty prompts in this dataset are previous translations from smugri-mt-bench and are included either verbatim or with minor modifications.

Prompt Structure

All prompts are multi-turn. A second turn was written for each prompt to introduce additional constraints, provide new information, or challenge the model to review its response.

Intended Use

The dataset is designed for human evaluation of model-generated responses. Tasks were selected to align with a predefined evaluation rubric.

Evaluation Protocol

Evaluators compare responses from two anonymous models and select a preferred response. If two responses are of equal quality, a tie may be selected.

Responses are evaluated along the following dimensions:

  • —Naturalness Which model's response sounds more natural in Estonian?
  • —Grammatical correctness Which model's response is grammatically more correct in Estonian?
  • —Sequential instruction following Which model is better at handling sequential instructions?
  • —Usefulness (applicable to writing and planning tasks only) Which model's response is more useful?
  • —Correctness (applicable to mathematics and reasoning tasks only) Which model's response is mathematically or logically more correct?
  • —Overall preference Which model's response do you prefer overall?