CoolFace
Datasetpublic

AgentPublic/evalap-meta-llama_llama-4-scout-17b-16e-instruct_mfs-32

meta-llama_Llama-4-Scout-17B-16E-Instruct_mfs (ID: 32) Experiment set for meta-llama_Llama-4-Scout-17B-16E-Instruct_mfs Overview This dataset contains 8 experiments from the EvalAP evaluation platform. Datasets: MFS_questions_v01 Models evaluated: meta-llama/Llama-4-Scout-17B-16E-Instruct Metrics: answer_relevancy, generation_time, judge_exactness, judge_notator, output_length Scores MFS_questions_v01 model answer_relevancy… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-meta-llama_llama-4-scout-17b-16e-instruct_mfs-32.

sourceHugging Faceupdated 8mo agoView on Hugging Face
0likes27downloads
Dataset Card

meta-llamaLlama-4-Scout-17B-16E-Instructmfs (ID: 32)

Experiment set for meta-llamaLlama-4-Scout-17B-16E-Instructmfs

Overview

This dataset contains 8 experiments from the EvalAP evaluation platform.

Datasets: MFSquestionsv01

Models evaluated: meta-llama/Llama-4-Scout-17B-16E-Instruct

Metrics: answerrelevancy, generationtime, judgeexactness, judgenotator, output_length

Scores

MFSquestionsv01

modelanswer_relevancygeneration_timejudge_exactnessjudge_notatoroutput_length
meta-llama/Llama-4-Scout-17B-16E-Instruct0.85 ± 0.216.97 ± 4.630.06 ± 0.235.50 ± 2.17293.34 ± 185.03

Usage

Use the dropdown above to select an experiment configuration.