CoolFace
Datasetpublic

AgentPublic/evalap-comparing-openweight-with-openaigpt-oss-120b-113

Comparing openweight with openai/gpt-oss-120b (ID: 113) Comparing openweight Albert-API with specific judge openai/gpt-oss-120b Overview This dataset contains 24 experiments from the EvalAP evaluation platform. Datasets: Assistant IA - QA, MFS_questions_v01 Models evaluated: openweight-large, openweight-medium, openweight-small Metrics: energy_consumption, generation_time, gwp_consumption, judge_notator, judge_precision, nb_tokens_completion, nb_tokens_prompt… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-comparing-openweight-with-openaigpt-oss-120b-113.

sourceHugging Faceupdated 8mo agoView on Hugging Face
0likes34downloads
Dataset Card

Comparing openweight with openai/gpt-oss-120b (ID: 113)

Comparing openweight Albert-API with specific judge openai/gpt-oss-120b

Overview

This dataset contains 24 experiments from the EvalAP evaluation platform.

Datasets: Assistant IA - QA, MFSquestionsv01

Models evaluated: openweight-large, openweight-medium, openweight-small

Metrics: energyconsumption, generationtime, gwpconsumption, judgenotator, judgeprecision, nbtokenscompletion, nbtokens_prompt

Scores

MFSquestionsv01

modelenergy_consumptiongeneration_timegwp_consumptionjudge_notatorjudge_precisionnb_tokens_completionnb_tokens_prompt
openai/gpt-oss-120b Albert0.02 ± 0.0114.46 ± 4.560.00 ± 0.006.70 ± 3.070.26 ± 0.441747.21 ± 615.2138.90 ± 9.26
Mistral small0.00 ± 0.0010.97 ± 5.620.00 ± 0.006.54 ± 2.840.16 ± 0.37456.58 ± 165.1038.90 ± 9.26
Qwen/Qwen3-VL-8B-Thinking0.00 ± 0.008.53 ± 19.810.00 ± 0.006.96 ± 2.820.21 ± 0.401365.38 ± 613.2038.90 ± 9.26

Assistant IA - QA

modelenergy_consumptiongeneration_timegwp_consumptionjudge_notatorjudge_precisionnb_tokens_completionnb_tokens_prompt
openai/gpt-oss-120b Albert0.01 ± 0.0110.23 ± 4.940.00 ± 0.006.41 ± 3.360.24 ± 0.431194.73 ± 714.0329.59 ± 7.05
Mistral small0.00 ± 0.008.54 ± 4.920.00 ± 0.006.04 ± 3.200.11 ± 0.32326.14 ± 139.1629.59 ± 7.05
Qwen/Qwen3-VL-8B-Thinking0.00 ± 0.004.90 ± 3.350.00 ± 0.005.58 ± 3.520.19 ± 0.391024.26 ± 692.9429.59 ± 7.05

Usage

Use the dropdown above to select an experiment configuration.