datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
evalap-comparing-openweight-with-patronusaiglider-114
Comparing openweight with PatronusAI/glider (ID: 114)
Comparing openweight Albert-API with specific judge PatronusAI/glider
Overview
This dataset contains 24 experiments
from the EvalAP evaluation platform.
Datasets: Assistant IA - QA, MFS_questions_v01
Models evaluated: openweight-large, openweight-medium, openweight-small
Metrics: energy_consumption, generation_time, gwp_consumption, judge_notator, judge_precision, nb_tokens_completion, nb_tokens_prompt
Scores… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-comparing-openweight-with-patronusaiglider-114.evalap-comparing-albert-api-models-v11-12-2025-107
Comparing Albert-API models v11-12-2025 (ID: 107)
Comparing albert models on MFS-AIA datasets
Overview
This dataset contains 20 experiments
from the EvalAP evaluation platform.
Datasets: Assistant IA - QA, MFS_questions_v01
Models evaluated: albert-large, albert-small, openweight-large, openweight-medium, openweight-small
Metrics: generation_time, judge_notator, judge_precision, nb_tokens_completion, nb_tokens_prompt, output_length
Scores
Assistant IA… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-comparing-albert-api-models-v11-12-2025-107.evalap-legalbenchrag-evaluation-v1-111
LegalBenchRAG Evaluation v1 (ID: 111)
A extensive RAG evaluation on the LegalBenchRAG dataset. See [complete me]
Overview
This dataset contains 36 experiments
from the EvalAP evaluation platform.
Datasets: LegalBenchRAG
Models evaluated: mistralai/Mistral-Small-3.2-24B-Instruct-2506
Metrics: judge_precision, output_length
Scores
LegalBenchRAG
model
judge_precision
output_length
model_semantic_20_qwen3_lbrv5
0.82 ± 0.38
163.13 ± 157.94… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-legalbenchrag-evaluation-v1-111.evalap-mfs_vllm_arena_v2-21
mfs_vllm_arena_v2 (ID: 21)
Experiment set for mfs_vllm_arena
Overview
This dataset contains 41 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: google/gemma-3-27b-it, meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503, neuralmagic/Meta-Llama-3.1-70B-Instruct-FP8
Metrics: answer_relevancy, generation_time, judge_exactness, judge_notator, output_length
Scores… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_vllm_arena_v2-21.evalap-mfs_tooling_v6-49
mfs_tooling_v6 (ID: 49)
Evaluating tooling capabilities.embedding model: bge-multilingual-gemma2collections: chunks-v13-04-25, limit=8
Overview
This dataset contains 24 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503
Metrics: generation_time, judge_precision, nb_tool_calls, output_length
Scores
MFS_questions_v01
model… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_tooling_v6-49.evalap-mfs_variability_v2-10
mfs_variability_v2 (ID: 10)
Comparing some models variability.
Overview
This dataset contains 70 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: AgentPublic/llama3-instruct-guillaumetell, google/gemma-2-9b-it, meta-llama/Llama-3.1-8B-Instruct, meta-llama/Llama-3.2-3B-Instruct, meta-llama/Llama-3.3-70B-Instruct, neuralmagic/Meta-Llama-3.1-70B-Instruct-FP8
Metrics: answer_relevancy, generation_time, judge_exactness… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_variability_v2-10.evalap-albert-api-rag-mfs-v3-68
albert-api-rag-mfs-v3 (ID: 68)
Evaluating hybrid search on MFS dataset.
Overview
This dataset contains 40 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: mistralai/Mistral-Small-3.1-24B-Instruct-2503, mistralai/Mistral-Small-3.2-24B-Instruct-2506
Metrics: generation_time, judge_precision, output_length
Scores
MFS_questions_v01
model
generation_time
judge_precision
output_length… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-albert-api-rag-mfs-v3-68.evalap-assistant-lasuite-tools-comparison-v2-101
Assistant LASuite Tools Comparison v2 (ID: 101)
Generated locally via notebook and pushed to EvalAP
Overview
This dataset contains 13 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Metrics: answer_relevancy, faithfulness, judge_exactness, judge_notator, judge_precision
Scores
MFS_questions_v01
model
answer_relevancy
judge_exactness
judge_notator
judge_precision
Mistral Medium (With WEB)(no sources)
0.95 ±… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-assistant-lasuite-tools-comparison-v2-101.evalap-mfs_tooling_v5-48
mfs_tooling_v5 (ID: 48)
Evaluating tooling capabilities.
Overview
This dataset contains 24 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503
Metrics: generation_time, judge_precision, nb_tool_calls, output_length
Scores
MFS_questions_v01
model
generation_time
judge_precision
nb_tool_calls
output_length… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_tooling_v5-48.evalap-gemma3_variability-62
gemma3_variability (ID: 62)
Comparing Gemma 3 4b variability across multiple temperatures.
Overview
This dataset contains 60 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: google/gemma-3-4b-it
Metrics: answer_relevancy, generation_time, judge_exactness, judge_notator, output_length
Scores
MFS_questions_v01
model
google/gemma-3-4b-it
Usage
Use the dropdown above to select an… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-gemma3_variability-62.evalap-compare-open-weight-models-31th-83
Compare Open Weight Models 31th (ID: 83)
Comparing open weight models
Overview
This dataset contains 44 experiments
from the EvalAP evaluation platform.
Datasets: Assistant IA - QA, MFS_questions_v01
Models evaluated: Groq/Llama-3-Groq-8B-Tool-Use, Qwen/Qwen3-VL-4B-Instruct, Qwen/Qwen3-VL-8B-Thinking, meta-llama/Llama-3.1-8B-Instruct, mistral-medium-2508, mistralai/Magistral-Small-2509, mistralai/Mistral-Small-3.2-24B-Instruct-2506… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-compare-open-weight-models-31th-83.evalap-comparing-mistral-medium-on-piaf-109
Comparing Mistral-Medium on piaf (ID: 109)
Comparing Mistral-Medium in different instances on piaf dataset
Overview
This dataset contains 6 experiments
from the EvalAP evaluation platform.
Datasets: piaf-v1.2
Models evaluated: mistral-medium-2508
Metrics: energy_consumption, generation_time, gwp_consumption, judge_notator, judge_precision, nb_tokens_completion, nb_tokens_prompt
Scores
piaf-v1.2
model
energy_consumption
generation_time… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-comparing-mistral-medium-on-piaf-109.evalap-comparing-openweight-with-openaigpt-oss-120b-113
Comparing openweight with openai/gpt-oss-120b (ID: 113)
Comparing openweight Albert-API with specific judge openai/gpt-oss-120b
Overview
This dataset contains 24 experiments
from the EvalAP evaluation platform.
Datasets: Assistant IA - QA, MFS_questions_v01
Models evaluated: openweight-large, openweight-medium, openweight-small
Metrics: energy_consumption, generation_time, gwp_consumption, judge_notator, judge_precision, nb_tokens_completion, nb_tokens_prompt… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-comparing-openweight-with-openaigpt-oss-120b-113.evalap-mediatechs-legi-chunking-evaluation-v1-116
MediaTech's LEGI Chunking Evaluation V1 (ID: 116)
Evaluation of severals chunking strategies for MediaTech's LEGI dataset.
Overview
This dataset contains 51 experiments
from the EvalAP evaluation platform.
Datasets: LEGI Synthetic QA Dataset
Metrics: contextual_precision, contextual_recall, contextual_relevancy, faithfulness, judge_precision
Scores
LEGI Synthetic QA Dataset
model
contextual_precision
contextual_recall
contextual_relevancy… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mediatechs-legi-chunking-evaluation-v1-116.evalap-mfs_variability_prompt_v1-22
mfs_variability_prompt_v1 (ID: 22)
Comparing impact of prompt system.
Overview
This dataset contains 24 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: gpt-3.5-turbo, meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503
Metrics: answer_relevancy, generation_time, judge_exactness, judge_notator, output_length
Scores
MFS_questions_v01
model
answer_relevancy
generation_time… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_variability_prompt_v1-22.evalap-comparing-albert-api-models-v11-12-2025-with-sysprompt-108
Comparing Albert-API models v11-12-2025 (with sysprompt) (ID: 108)
Comparing albert models on MFS-AIA datasets (with sysprompt)
Overview
This dataset contains 20 experiments
from the EvalAP evaluation platform.
Datasets: Assistant IA - QA, MFS_questions_v01
Models evaluated: albert-large, albert-small, openweight-large, openweight-medium, openweight-small
Metrics: generation_time, judge_notator, judge_precision, nb_tokens_completion, nb_tokens_prompt, output_length… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-comparing-albert-api-models-v11-12-2025-with-sysprompt-108.evalap-albert_small_variability-63
albert_small_variability (ID: 63)
Comparing Albert small variability across multiple temperatures.
Overview
This dataset contains 60 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: meta-llama/Llama-3.1-8B-Instruct
Metrics: answer_relevancy, generation_time, judge_exactness, judge_notator, output_length
Scores
MFS_questions_v01
model
answer_relevancy
generation_time
judge_exactness
judge_notator… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-albert_small_variability-63.evalap-mfs_tooling_v4-47
mfs_tooling_v4 (ID: 47)
Evaluating tooling capabilities.
Overview
This dataset contains 24 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503
Metrics: contextual_relevancy, generation_time, judge_exactness, judge_notator, judge_precision, nb_tool_calls, output_length
Scores
MFS_questions_v01
model
contextual_relevancy… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_tooling_v4-47.evalap-mfs_tooling_v1-33
mfs_tooling_v1 (ID: 33)
Evaluating tooling capabilities.
Overview
This dataset contains 26 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503
Metrics: contextual_relevancy, faithfulness, generation_time, judge_exactness, judge_notator, nb_tool_calls, output_length, ragas
Scores
MFS_questions_v01
model
contextual_relevancy… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_tooling_v1-33.evalap-mfs_with_judge_deepseek-aideepseek-r1-distill-qwen-32b_v5-52
mfs_with_judge_deepseek-ai/DeepSeek-R1-Distill-Qwen-32B_v5 (ID: 52)
Comparing impact of judge in score calculation. Here : deepseek-ai/DeepSeek-R1-Distill-Qwen-32B
Overview
This dataset contains 19 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: google/gemma-2-9b-it, gpt-3.5-turbo, meta-llama/Llama-3.1-8B-Instruct, meta-llama/Llama-3.3-70B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_with_judge_deepseek-aideepseek-r1-distill-qwen-32b_v5-52.evalap-mfs_rag_limit_v1-7
mfs_rag_limit_v1 (ID: 7)
Comparing the impact of the limit parameters on a RAG model.
Overview
This dataset contains 9 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: AgentPublic/llama3-instruct-8b
Metrics: answer_relevancy, generation_time, judge_exactness, judge_notator, output_length
Scores
MFS_questions_v01
model
answer_relevancy
generation_time
judge_exactness
judge_notator
output_length… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_rag_limit_v1-7.evalap-meta-llama_llama-4-scout-17b-16e-instruct_mfs-32
meta-llama_Llama-4-Scout-17B-16E-Instruct_mfs (ID: 32)
Experiment set for meta-llama_Llama-4-Scout-17B-16E-Instruct_mfs
Overview
This dataset contains 8 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: meta-llama/Llama-4-Scout-17B-16E-Instruct
Metrics: answer_relevancy, generation_time, judge_exactness, judge_notator, output_length
Scores
MFS_questions_v01
model
answer_relevancy
generation_time… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-meta-llama_llama-4-scout-17b-16e-instruct_mfs-32.evalap-mfs_tooling_v7-50
mfs_tooling_v7 (ID: 50)
Evaluating tooling capabilities.embedding model: BAAI/bge-m3collections: chunks-v6, limit=10
Overview
This dataset contains 24 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503
Metrics: generation_time, judge_precision, nb_tool_calls, output_length
Scores
MFS_questions_v01
model
generation_time… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_tooling_v7-50.evalap-albert-api-rag-mfs-v1-65
albert-api-rag-mfs-v1 (ID: 65)
Evaluating hybrid search on MFS dataset.
Overview
This dataset contains 20 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: mistralai/Mistral-Small-3.1-24B-Instruct-2503
Metrics: generation_time, judge_precision, output_length
Scores
MFS_questions_v01
model
generation_time
judge_precision
output_length
albert-large
14.77 ± 5.35
0.30 ± 0.46
363.56 ± 105.75… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-albert-api-rag-mfs-v1-65.evalap-mfs_variability_prompt_v2-23
mfs_variability_prompt_v2 (ID: 23)
Comparing impact of prompt system with gpt-4o in judge.
Overview
This dataset contains 27 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: gpt-3.5-turbo, meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503
Metrics: answer_relevancy, generation_time, judge_exactness, judge_notator, output_length
Scores
MFS_questions_v01
model… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_variability_prompt_v2-23.evalap-mfs_tooling_v3-44
mfs_tooling_v3 (ID: 44)
Evaluating tooling capabilities.
Overview
This dataset contains 24 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503
Metrics: contextual_relevancy, faithfulness, generation_time, judge_exactness, judge_notator, nb_tool_calls, output_length
Scores
MFS_questions_v01
model
contextual_relevancy… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_tooling_v3-44.evalap-mfs_tooling_v8-51
mfs_tooling_v8 (ID: 51)
Evaluating tooling capabilities.embedding model: bge-multilingual-gemma2collections: chunks-v13-04-25, limit=10
Overview
This dataset contains 24 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503
Metrics: generation_time, judge_precision, nb_tool_calls, output_length
Scores
MFS_questions_v01
model… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_tooling_v8-51.evalap-baseline_model_assia-67
baseline_model_assIA (ID: 67)
Baseline evaluation for Assistant IA dataset
Overview
This dataset contains 20 experiments
from the EvalAP evaluation platform.
Datasets: Assistant IA - QA
Models evaluated: albert-large, albert-small, openai/gpt-oss-120b, openai/gpt-oss-20b
Metrics: energy_consumption, generation_time, gwp_consumption, judge_notator, judge_precision, nb_tokens_completion, nb_tokens_prompt, nb_tool_calls
Scores
Assistant IA - QA… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-baseline_model_assia-67.evalap-deccp-v1-79
DECCP v1 (ID: 79)
DECCP Evaluation with Lllm-as-a-Judge
Overview
This dataset contains 8 experiments
from the EvalAP evaluation platform.
Datasets: DECCP
Models evaluated: Qwen/Qwen3-VL-8B-Thinking, deepseek-r1-0528, meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.2-24B-Instruct-2506
Metrics: answer_relevancy, judge_censorship
Scores
DECCP
model
answer_relevancy
judge_censorship
Qwen/Qwen3-VL-8B-Thinking
0.32 ± 0.31
0.32 ±… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-deccp-v1-79.evalap-wikipedia_frames_150-11
Wikipedia_Frames_150 (ID: 11)
Compare DeepSearch on Rag, vannilla models on complex dataset.
Overview
This dataset contains 6 experiments
from the EvalAP evaluation platform.
Datasets: WikipediaFrames_150
Metrics: answer_relevancy, judge_exactness, judge_notator, output_length
Scores
WikipediaFrames_150
model
answer_relevancy
judge_exactness
judge_notator
output_length
deepsearch_8B(3.1)70B(3.3)-web_3_3_3
0.73 ± 0.44
0.27 ± 0.44
3.62 ±… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-wikipedia_frames_150-11.
