CoolFace
Datasetpublic

AgentPublic/evalap-comparing-albert-api-models-v11-12-2025-with-sysprompt-108

Comparing Albert-API models v11-12-2025 (with sysprompt) (ID: 108) Comparing albert models on MFS-AIA datasets (with sysprompt) Overview This dataset contains 20 experiments from the EvalAP evaluation platform. Datasets: Assistant IA - QA, MFS_questions_v01 Models evaluated: albert-large, albert-small, openweight-large, openweight-medium, openweight-small Metrics: generation_time, judge_notator, judge_precision, nb_tokens_completion, nb_tokens_prompt… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-comparing-albert-api-models-v11-12-2025-with-sysprompt-108.

sourceHugging Faceupdated 8mo agoView on Hugging Face
0likes35downloads
Dataset Card

Comparing Albert-API models v11-12-2025 (with sysprompt) (ID: 108)

Comparing albert models on MFS-AIA datasets (with sysprompt)

Overview

This dataset contains 20 experiments from the EvalAP evaluation platform.

Datasets: Assistant IA - QA, MFSquestionsv01

Models evaluated: albert-large, albert-small, openweight-large, openweight-medium, openweight-small

Metrics: generationtime, judgenotator, judgeprecision, nbtokenscompletion, nbtokensprompt, outputlength

Scores

Assistant IA - QA

modelgeneration_timejudge_notatorjudge_precisionnb_tokens_completionnb_tokens_promptoutput_length
albert-large0.93 ± 0.925.33 ± 2.890.03 ± 0.1828.54 ± 36.8728.59 ± 7.0518.74 ± 25.54
albert-small3.73 ± 25.953.55 ± 3.030.10 ± 0.3051.70 ± 66.1328.59 ± 7.0534.30 ± 43.20
openweight-large1.68 ± 0.816.14 ± 3.110.38 ± 0.4960.11 ± 72.6828.59 ± 7.0537.82 ± 45.77
openweight-medium0.45 ± 0.675.45 ± 2.840.03 ± 0.1827.46 ± 34.9428.59 ± 7.0517.93 ± 24.15
openweight-small0.29 ± 0.525.09 ± 3.010.14 ± 0.3560.62 ± 79.0928.59 ± 7.0530.55 ± 40.17

MFSquestionsv01

modelgeneration_timejudge_notatorjudge_precisionnb_tokens_completionnb_tokens_promptoutput_length
albert-large1.35 ± 1.116.47 ± 2.090.17 ± 0.3741.45 ± 36.5837.90 ± 9.2627.72 ± 23.56
albert-small14.03 ± 54.244.78 ± 2.520.13 ± 0.3370.76 ± 66.2437.90 ± 9.2650.38 ± 47.97
openweight-large1.74 ± 0.786.86 ± 2.260.46 ± 0.5095.62 ± 72.8537.90 ± 9.2656.85 ± 41.96
openweight-medium0.69 ± 0.696.44 ± 2.230.18 ± 0.3839.32 ± 36.8337.90 ± 9.2626.22 ± 23.99
openweight-small0.53 ± 0.506.67 ± 2.180.32 ± 0.4777.13 ± 53.7237.90 ± 9.2638.13 ± 27.00

Usage

Use the dropdown above to select an experiment configuration.