AgentPublic/evalap-mfs_variability_v2-10
mfs_variability_v2 (ID: 10) Comparing some models variability. Overview This dataset contains 70 experiments from the EvalAP evaluation platform. Datasets: MFS_questions_v01 Models evaluated: AgentPublic/llama3-instruct-guillaumetell, google/gemma-2-9b-it, meta-llama/Llama-3.1-8B-Instruct, meta-llama/Llama-3.2-3B-Instruct, meta-llama/Llama-3.3-70B-Instruct, neuralmagic/Meta-Llama-3.1-70B-Instruct-FP8 Metrics: answer_relevancy, generation_time, judge_exactness… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_variability_v2-10.
mfsvariabilityv2 (ID: 10)
Comparing some models variability.
Overview
This dataset contains 70 experiments from the EvalAP evaluation platform.
Datasets: MFSquestionsv01
Models evaluated: AgentPublic/llama3-instruct-guillaumetell, google/gemma-2-9b-it, meta-llama/Llama-3.1-8B-Instruct, meta-llama/Llama-3.2-3B-Instruct, meta-llama/Llama-3.3-70B-Instruct, neuralmagic/Meta-Llama-3.1-70B-Instruct-FP8
Metrics: answerrelevancy, generationtime, judgeexactness, judgenotator, output_length
Scores
MFSquestionsv01
Usage
Use the dropdown above to select an experiment configuration.
