CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AgentPublic /evalap-comparing-openweight-with-patronusaiglider-114 Comparing openweight with PatronusAI/glider (ID: 114) Comparing openweight Albert-API with specific judge PatronusAI/glider Overview This dataset contains 24 experiments from the EvalAP evaluation platform. Datasets: Assistant IA - QA, MFS_questions_v01 Models evaluated: openweight-large, openweight-medium, openweight-small Metrics: energy_consumption, generation_time, gwp_consumption, judge_notator, judge_precision, nb_tokens_completion, nb_tokens_prompt Scores… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-comparing-openweight-with-patronusaiglider-114.tabularn<1K0 likes80 downloads8mo agoHugging Face02AgentPublic /evalap-comparing-albert-api-models-v11-12-2025-107 Comparing Albert-API models v11-12-2025 (ID: 107) Comparing albert models on MFS-AIA datasets Overview This dataset contains 20 experiments from the EvalAP evaluation platform. Datasets: Assistant IA - QA, MFS_questions_v01 Models evaluated: albert-large, albert-small, openweight-large, openweight-medium, openweight-small Metrics: generation_time, judge_notator, judge_precision, nb_tokens_completion, nb_tokens_prompt, output_length Scores Assistant IA… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-comparing-albert-api-models-v11-12-2025-107.tabularn<1K0 likes75 downloads8mo agoHugging Face03AgentPublic /evalap-legalbenchrag-evaluation-v1-111 LegalBenchRAG Evaluation v1 (ID: 111) A extensive RAG evaluation on the LegalBenchRAG dataset. See [complete me] Overview This dataset contains 36 experiments from the EvalAP evaluation platform. Datasets: LegalBenchRAG Models evaluated: mistralai/Mistral-Small-3.2-24B-Instruct-2506 Metrics: judge_precision, output_length Scores LegalBenchRAG model judge_precision output_length model_semantic_20_qwen3_lbrv5 0.82 ± 0.38 163.13 ± 157.94… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-legalbenchrag-evaluation-v1-111.tabular10K<n<100K0 likes54 downloads8mo agoHugging Face04AgentPublic /evalap-mfs_vllm_arena_v2-21 mfs_vllm_arena_v2 (ID: 21) Experiment set for mfs_vllm_arena Overview This dataset contains 41 experiments from the EvalAP evaluation platform. Datasets: MFS_questions_v01 Models evaluated: google/gemma-3-27b-it, meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503, neuralmagic/Meta-Llama-3.1-70B-Instruct-FP8 Metrics: answer_relevancy, generation_time, judge_exactness, judge_notator, output_length Scores… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_vllm_arena_v2-21.tabular1K<n<10K0 likes48 downloads8mo agoHugging Face05AgentPublic /evalap-mfs_tooling_v6-49 mfs_tooling_v6 (ID: 49) Evaluating tooling capabilities.embedding model: bge-multilingual-gemma2collections: chunks-v13-04-25, limit=8 Overview This dataset contains 24 experiments from the EvalAP evaluation platform. Datasets: MFS_questions_v01 Models evaluated: meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503 Metrics: generation_time, judge_precision, nb_tool_calls, output_length Scores MFS_questions_v01 model… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_tooling_v6-49.tabularn<1K0 likes48 downloads8mo agoHugging Face06AgentPublic /evalap-mfs_variability_v2-10 mfs_variability_v2 (ID: 10) Comparing some models variability. Overview This dataset contains 70 experiments from the EvalAP evaluation platform. Datasets: MFS_questions_v01 Models evaluated: AgentPublic/llama3-instruct-guillaumetell, google/gemma-2-9b-it, meta-llama/Llama-3.1-8B-Instruct, meta-llama/Llama-3.2-3B-Instruct, meta-llama/Llama-3.3-70B-Instruct, neuralmagic/Meta-Llama-3.1-70B-Instruct-FP8 Metrics: answer_relevancy, generation_time, judge_exactness… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_variability_v2-10.tabular1K<n<10K0 likes46 downloads8mo agoHugging Face07AgentPublic /evalap-albert-api-rag-mfs-v3-68 albert-api-rag-mfs-v3 (ID: 68) Evaluating hybrid search on MFS dataset. Overview This dataset contains 40 experiments from the EvalAP evaluation platform. Datasets: MFS_questions_v01 Models evaluated: mistralai/Mistral-Small-3.1-24B-Instruct-2503, mistralai/Mistral-Small-3.2-24B-Instruct-2506 Metrics: generation_time, judge_precision, output_length Scores MFS_questions_v01 model generation_time judge_precision output_length… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-albert-api-rag-mfs-v3-68.tabular1K<n<10K0 likes46 downloads8mo agoHugging Face08AgentPublic /evalap-assistant-lasuite-tools-comparison-v2-101 Assistant LASuite Tools Comparison v2 (ID: 101) Generated locally via notebook and pushed to EvalAP Overview This dataset contains 13 experiments from the EvalAP evaluation platform. Datasets: MFS_questions_v01 Metrics: answer_relevancy, faithfulness, judge_exactness, judge_notator, judge_precision Scores MFS_questions_v01 model answer_relevancy judge_exactness judge_notator judge_precision Mistral Medium (With WEB)(no sources) 0.95 ±… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-assistant-lasuite-tools-comparison-v2-101.tabularn<1K0 likes46 downloads8mo agoHugging Face09AgentPublic /evalap-mfs_tooling_v5-48 mfs_tooling_v5 (ID: 48) Evaluating tooling capabilities. Overview This dataset contains 24 experiments from the EvalAP evaluation platform. Datasets: MFS_questions_v01 Models evaluated: meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503 Metrics: generation_time, judge_precision, nb_tool_calls, output_length Scores MFS_questions_v01 model generation_time judge_precision nb_tool_calls output_length… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_tooling_v5-48.tabularn<1K0 likes42 downloads8mo agoHugging Face10AgentPublic /evalap-gemma3_variability-62 gemma3_variability (ID: 62) Comparing Gemma 3 4b variability across multiple temperatures. Overview This dataset contains 60 experiments from the EvalAP evaluation platform. Datasets: MFS_questions_v01 Models evaluated: google/gemma-3-4b-it Metrics: answer_relevancy, generation_time, judge_exactness, judge_notator, output_length Scores MFS_questions_v01 model google/gemma-3-4b-it Usage Use the dropdown above to select an… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-gemma3_variability-62.tabular1K<n<10K0 likes41 downloads8mo agoHugging Face11AgentPublic /evalap-compare-open-weight-models-31th-83 Compare Open Weight Models 31th (ID: 83) Comparing open weight models Overview This dataset contains 44 experiments from the EvalAP evaluation platform. Datasets: Assistant IA - QA, MFS_questions_v01 Models evaluated: Groq/Llama-3-Groq-8B-Tool-Use, Qwen/Qwen3-VL-4B-Instruct, Qwen/Qwen3-VL-8B-Thinking, meta-llama/Llama-3.1-8B-Instruct, mistral-medium-2508, mistralai/Magistral-Small-2509, mistralai/Mistral-Small-3.2-24B-Instruct-2506… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-compare-open-weight-models-31th-83.tabular1K<n<10K0 likes39 downloads8mo agoHugging Face12AgentPublic /evalap-comparing-mistral-medium-on-piaf-109 Comparing Mistral-Medium on piaf (ID: 109) Comparing Mistral-Medium in different instances on piaf dataset Overview This dataset contains 6 experiments from the EvalAP evaluation platform. Datasets: piaf-v1.2 Models evaluated: mistral-medium-2508 Metrics: energy_consumption, generation_time, gwp_consumption, judge_notator, judge_precision, nb_tokens_completion, nb_tokens_prompt Scores piaf-v1.2 model energy_consumption generation_time… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-comparing-mistral-medium-on-piaf-109.tabular10K<n<100K0 likes38 downloads8mo agoHugging Face13AgentPublic /evalap-comparing-openweight-with-openaigpt-oss-120b-113 Comparing openweight with openai/gpt-oss-120b (ID: 113) Comparing openweight Albert-API with specific judge openai/gpt-oss-120b Overview This dataset contains 24 experiments from the EvalAP evaluation platform. Datasets: Assistant IA - QA, MFS_questions_v01 Models evaluated: openweight-large, openweight-medium, openweight-small Metrics: energy_consumption, generation_time, gwp_consumption, judge_notator, judge_precision, nb_tokens_completion, nb_tokens_prompt… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-comparing-openweight-with-openaigpt-oss-120b-113.tabular1K<n<10K0 likes37 downloads8mo agoHugging Face14AgentPublic /evalap-mediatechs-legi-chunking-evaluation-v1-116 MediaTech's LEGI Chunking Evaluation V1 (ID: 116) Evaluation of severals chunking strategies for MediaTech's LEGI dataset. Overview This dataset contains 51 experiments from the EvalAP evaluation platform. Datasets: LEGI Synthetic QA Dataset Metrics: contextual_precision, contextual_recall, contextual_relevancy, faithfulness, judge_precision Scores LEGI Synthetic QA Dataset model contextual_precision contextual_recall contextual_relevancy… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mediatechs-legi-chunking-evaluation-v1-116.tabular1K<n<10K0 likes37 downloads8mo agoHugging Face15AgentPublic /evalap-mfs_variability_prompt_v1-22 mfs_variability_prompt_v1 (ID: 22) Comparing impact of prompt system. Overview This dataset contains 24 experiments from the EvalAP evaluation platform. Datasets: MFS_questions_v01 Models evaluated: gpt-3.5-turbo, meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503 Metrics: answer_relevancy, generation_time, judge_exactness, judge_notator, output_length Scores MFS_questions_v01 model answer_relevancy generation_time… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_variability_prompt_v1-22.tabularn<1K0 likes36 downloads8mo agoHugging Face16AgentPublic /evalap-comparing-albert-api-models-v11-12-2025-with-sysprompt-108 Comparing Albert-API models v11-12-2025 (with sysprompt) (ID: 108) Comparing albert models on MFS-AIA datasets (with sysprompt) Overview This dataset contains 20 experiments from the EvalAP evaluation platform. Datasets: Assistant IA - QA, MFS_questions_v01 Models evaluated: albert-large, albert-small, openweight-large, openweight-medium, openweight-small Metrics: generation_time, judge_notator, judge_precision, nb_tokens_completion, nb_tokens_prompt, output_length… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-comparing-albert-api-models-v11-12-2025-with-sysprompt-108.tabularn<1K0 likes35 downloads8mo agoHugging Face17AgentPublic /evalap-albert_small_variability-63 albert_small_variability (ID: 63) Comparing Albert small variability across multiple temperatures. Overview This dataset contains 60 experiments from the EvalAP evaluation platform. Datasets: MFS_questions_v01 Models evaluated: meta-llama/Llama-3.1-8B-Instruct Metrics: answer_relevancy, generation_time, judge_exactness, judge_notator, output_length Scores MFS_questions_v01 model answer_relevancy generation_time judge_exactness judge_notator… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-albert_small_variability-63.tabular1K<n<10K0 likes32 downloads8mo agoHugging Face18AgentPublic /evalap-mfs_tooling_v4-47 mfs_tooling_v4 (ID: 47) Evaluating tooling capabilities. Overview This dataset contains 24 experiments from the EvalAP evaluation platform. Datasets: MFS_questions_v01 Models evaluated: meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503 Metrics: contextual_relevancy, generation_time, judge_exactness, judge_notator, judge_precision, nb_tool_calls, output_length Scores MFS_questions_v01 model contextual_relevancy… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_tooling_v4-47.tabularn<1K0 likes31 downloads8mo agoHugging Face19AgentPublic /evalap-mfs_tooling_v1-33 mfs_tooling_v1 (ID: 33) Evaluating tooling capabilities. Overview This dataset contains 26 experiments from the EvalAP evaluation platform. Datasets: MFS_questions_v01 Models evaluated: meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503 Metrics: contextual_relevancy, faithfulness, generation_time, judge_exactness, judge_notator, nb_tool_calls, output_length, ragas Scores MFS_questions_v01 model contextual_relevancy… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_tooling_v1-33.tabular1K<n<10K0 likes30 downloads8mo agoHugging Face20AgentPublic /evalap-mfs_with_judge_deepseek-aideepseek-r1-distill-qwen-32b_v5-52 mfs_with_judge_deepseek-ai/DeepSeek-R1-Distill-Qwen-32B_v5 (ID: 52) Comparing impact of judge in score calculation. Here : deepseek-ai/DeepSeek-R1-Distill-Qwen-32B Overview This dataset contains 19 experiments from the EvalAP evaluation platform. Datasets: MFS_questions_v01 Models evaluated: google/gemma-2-9b-it, gpt-3.5-turbo, meta-llama/Llama-3.1-8B-Instruct, meta-llama/Llama-3.3-70B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_with_judge_deepseek-aideepseek-r1-distill-qwen-32b_v5-52.tabularn<1K0 likes30 downloads8mo agoHugging Face21AgentPublic /evalap-mfs_rag_limit_v1-7 mfs_rag_limit_v1 (ID: 7) Comparing the impact of the limit parameters on a RAG model. Overview This dataset contains 9 experiments from the EvalAP evaluation platform. Datasets: MFS_questions_v01 Models evaluated: AgentPublic/llama3-instruct-8b Metrics: answer_relevancy, generation_time, judge_exactness, judge_notator, output_length Scores MFS_questions_v01 model answer_relevancy generation_time judge_exactness judge_notator output_length… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_rag_limit_v1-7.tabularn<1K0 likes29 downloads8mo agoHugging Face22AgentPublic /evalap-meta-llama_llama-4-scout-17b-16e-instruct_mfs-32 meta-llama_Llama-4-Scout-17B-16E-Instruct_mfs (ID: 32) Experiment set for meta-llama_Llama-4-Scout-17B-16E-Instruct_mfs Overview This dataset contains 8 experiments from the EvalAP evaluation platform. Datasets: MFS_questions_v01 Models evaluated: meta-llama/Llama-4-Scout-17B-16E-Instruct Metrics: answer_relevancy, generation_time, judge_exactness, judge_notator, output_length Scores MFS_questions_v01 model answer_relevancy generation_time… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-meta-llama_llama-4-scout-17b-16e-instruct_mfs-32.tabularn<1K0 likes29 downloads8mo agoHugging Face23AgentPublic /evalap-mfs_tooling_v7-50 mfs_tooling_v7 (ID: 50) Evaluating tooling capabilities.embedding model: BAAI/bge-m3collections: chunks-v6, limit=10 Overview This dataset contains 24 experiments from the EvalAP evaluation platform. Datasets: MFS_questions_v01 Models evaluated: meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503 Metrics: generation_time, judge_precision, nb_tool_calls, output_length Scores MFS_questions_v01 model generation_time… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_tooling_v7-50.tabularn<1K0 likes29 downloads8mo agoHugging Face24AgentPublic /evalap-albert-api-rag-mfs-v1-65 albert-api-rag-mfs-v1 (ID: 65) Evaluating hybrid search on MFS dataset. Overview This dataset contains 20 experiments from the EvalAP evaluation platform. Datasets: MFS_questions_v01 Models evaluated: mistralai/Mistral-Small-3.1-24B-Instruct-2503 Metrics: generation_time, judge_precision, output_length Scores MFS_questions_v01 model generation_time judge_precision output_length albert-large 14.77 ± 5.35 0.30 ± 0.46 363.56 ± 105.75… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-albert-api-rag-mfs-v1-65.tabularn<1K0 likes28 downloads8mo agoHugging Face25AgentPublic /evalap-mfs_variability_prompt_v2-23 mfs_variability_prompt_v2 (ID: 23) Comparing impact of prompt system with gpt-4o in judge. Overview This dataset contains 27 experiments from the EvalAP evaluation platform. Datasets: MFS_questions_v01 Models evaluated: gpt-3.5-turbo, meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503 Metrics: answer_relevancy, generation_time, judge_exactness, judge_notator, output_length Scores MFS_questions_v01 model… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_variability_prompt_v2-23.tabular1K<n<10K0 likes27 downloads8mo agoHugging Face26AgentPublic /evalap-mfs_tooling_v3-44 mfs_tooling_v3 (ID: 44) Evaluating tooling capabilities. Overview This dataset contains 24 experiments from the EvalAP evaluation platform. Datasets: MFS_questions_v01 Models evaluated: meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503 Metrics: contextual_relevancy, faithfulness, generation_time, judge_exactness, judge_notator, nb_tool_calls, output_length Scores MFS_questions_v01 model contextual_relevancy… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_tooling_v3-44.tabularn<1K0 likes27 downloads8mo agoHugging Face27AgentPublic /evalap-mfs_tooling_v8-51 mfs_tooling_v8 (ID: 51) Evaluating tooling capabilities.embedding model: bge-multilingual-gemma2collections: chunks-v13-04-25, limit=10 Overview This dataset contains 24 experiments from the EvalAP evaluation platform. Datasets: MFS_questions_v01 Models evaluated: meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503 Metrics: generation_time, judge_precision, nb_tool_calls, output_length Scores MFS_questions_v01 model… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mfs_tooling_v8-51.tabularn<1K0 likes27 downloads8mo agoHugging Face28AgentPublic /evalap-baseline_model_assia-67 baseline_model_assIA (ID: 67) Baseline evaluation for Assistant IA dataset Overview This dataset contains 20 experiments from the EvalAP evaluation platform. Datasets: Assistant IA - QA Models evaluated: albert-large, albert-small, openai/gpt-oss-120b, openai/gpt-oss-20b Metrics: energy_consumption, generation_time, gwp_consumption, judge_notator, judge_precision, nb_tokens_completion, nb_tokens_prompt, nb_tool_calls Scores Assistant IA - QA… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-baseline_model_assia-67.tabularn<1K0 likes27 downloads8mo agoHugging Face29AgentPublic /evalap-deccp-v1-79 DECCP v1 (ID: 79) DECCP Evaluation with Lllm-as-a-Judge Overview This dataset contains 8 experiments from the EvalAP evaluation platform. Datasets: DECCP Models evaluated: Qwen/Qwen3-VL-8B-Thinking, deepseek-r1-0528, meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.2-24B-Instruct-2506 Metrics: answer_relevancy, judge_censorship Scores DECCP model answer_relevancy judge_censorship Qwen/Qwen3-VL-8B-Thinking 0.32 ± 0.31 0.32 ±… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-deccp-v1-79.tabularn<1K0 likes26 downloads8mo agoHugging Face30AgentPublic /evalap-wikipedia_frames_150-11 Wikipedia_Frames_150 (ID: 11) Compare DeepSearch on Rag, vannilla models on complex dataset. Overview This dataset contains 6 experiments from the EvalAP evaluation platform. Datasets: WikipediaFrames_150 Metrics: answer_relevancy, judge_exactness, judge_notator, output_length Scores WikipediaFrames_150 model answer_relevancy judge_exactness judge_notator output_length deepsearch_8B(3.1)70B(3.3)-web_3_3_3 0.73 ± 0.44 0.27 ± 0.44 3.62 ±… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-wikipedia_frames_150-11.tabularn<1K0 likes25 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.