CoolFace
Datasetpublic

AgentPublic/evalap-mfs_variability_v2-10

mfs_variability_v2 (ID: 10) Comparing some models variability. Overview This dataset contains 70 experiments from the EvalAP evaluation platform. Datasets: MFS_questions_v01 Models evaluated: AgentPublic/llama3-instruct-guillaumetell, google/gemma-2-9b-it, meta-llama/Llama-3.1-8B-Instruct, meta-llama/Llama-3.2-3B-Instruct, meta-llama/Llama-3.3-70B-Instruct, neuralmagic/Meta-Llama-3.1-70B-Instruct-FP8 Metrics: answer_relevancy, generation_time, judge_exactness… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_variability_v2-10.

sourceHugging Faceupdated 8mo agoView on Hugging Face
0likes46downloads
Dataset Card

mfsvariabilityv2 (ID: 10)

Comparing some models variability.

Overview

This dataset contains 70 experiments from the EvalAP evaluation platform.

Datasets: MFSquestionsv01

Models evaluated: AgentPublic/llama3-instruct-guillaumetell, google/gemma-2-9b-it, meta-llama/Llama-3.1-8B-Instruct, meta-llama/Llama-3.2-3B-Instruct, meta-llama/Llama-3.3-70B-Instruct, neuralmagic/Meta-Llama-3.1-70B-Instruct-FP8

Metrics: answerrelevancy, generationtime, judgeexactness, judgenotator, output_length

Scores

MFSquestionsv01

modelanswer_relevancygeneration_timejudge_exactnessjudge_notatoroutput_length
AgentPublic/llama3-instruct-guillaumetell0.78 ± 0.265.50 ± 3.290.13 ± 0.335.28 ± 2.64133.25 ± 69.29
meta-llama/Llama-3.1-8B-Instruct0.80 ± 0.237.51 ± 21.600.05 ± 0.224.27 ± 2.57212.47 ± 124.94
meta-llama/Llama-3.3-70B-Instruct0.85 ± 0.1912.27 ± 4.260.04 ± 0.194.34 ± 2.39339.05 ± 110.87
neuralmagic/Meta-Llama-3.1-70B-Instruct-FP80.89 ± 0.1616.82 ± 14.300.07 ± 0.254.77 ± 2.43266.96 ± 110.46
meta-llama/Llama-3.2-3B-Instruct0.74 ± 0.252.48 ± 0.990.00 ± 0.002.88 ± 1.71320.40 ± 121.72
google/gemma-2-9b-it0.77 ± 0.2231.30 ± 74.680.05 ± 0.214.37 ± 2.18229.90 ± 91.77

Usage

Use the dropdown above to select an experiment configuration.