evalap
Datasets
All datasets matching “evalap”evalap-comparing-openweight-with-patronusaiglider-114
Comparing openweight with PatronusAI/glider (ID: 114)
Comparing openweight Albert-API with specific judge PatronusAI/glider
Overview
This dataset contains 24 experiments
from the EvalAP evaluation platform.
Datasets: Assistant IA - QA, MFS_questions_v01
Models evaluated: openweight-large, openweight-medium, openweight-small
Metrics: energy_consumption, generation_time, gwp_consumption, judge_notator, judge_precision, nb_tokens_completion, nb_tokens_prompt
Scores… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-comparing-openweight-with-patronusaiglider-114.evalap-comparing-albert-api-models-v11-12-2025-107
Comparing Albert-API models v11-12-2025 (ID: 107)
Comparing albert models on MFS-AIA datasets
Overview
This dataset contains 20 experiments
from the EvalAP evaluation platform.
Datasets: Assistant IA - QA, MFS_questions_v01
Models evaluated: albert-large, albert-small, openweight-large, openweight-medium, openweight-small
Metrics: generation_time, judge_notator, judge_precision, nb_tokens_completion, nb_tokens_prompt, output_length
Scores
Assistant IA… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-comparing-albert-api-models-v11-12-2025-107.evalap-legalbenchrag-evaluation-v1-111
LegalBenchRAG Evaluation v1 (ID: 111)
A extensive RAG evaluation on the LegalBenchRAG dataset. See [complete me]
Overview
This dataset contains 36 experiments
from the EvalAP evaluation platform.
Datasets: LegalBenchRAG
Models evaluated: mistralai/Mistral-Small-3.2-24B-Instruct-2506
Metrics: judge_precision, output_length
Scores
LegalBenchRAG
model
judge_precision
output_length
model_semantic_20_qwen3_lbrv5
0.82 ± 0.38
163.13 ± 157.94… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-legalbenchrag-evaluation-v1-111.evalap-mfs_vllm_arena_v2-21
mfs_vllm_arena_v2 (ID: 21)
Experiment set for mfs_vllm_arena
Overview
This dataset contains 41 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: google/gemma-3-27b-it, meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503, neuralmagic/Meta-Llama-3.1-70B-Instruct-FP8
Metrics: answer_relevancy, generation_time, judge_exactness, judge_notator, output_length
Scores… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_vllm_arena_v2-21.evalap-mfs_tooling_v6-49
mfs_tooling_v6 (ID: 49)
Evaluating tooling capabilities.embedding model: bge-multilingual-gemma2collections: chunks-v13-04-25, limit=8
Overview
This dataset contains 24 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503
Metrics: generation_time, judge_precision, nb_tool_calls, output_length
Scores
MFS_questions_v01
model… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_tooling_v6-49.evalap-mfs_variability_v2-10
mfs_variability_v2 (ID: 10)
Comparing some models variability.
Overview
This dataset contains 70 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: AgentPublic/llama3-instruct-guillaumetell, google/gemma-2-9b-it, meta-llama/Llama-3.1-8B-Instruct, meta-llama/Llama-3.2-3B-Instruct, meta-llama/Llama-3.3-70B-Instruct, neuralmagic/Meta-Llama-3.1-70B-Instruct-FP8
Metrics: answer_relevancy, generation_time, judge_exactness… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_variability_v2-10.
