CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cjvt /slovenian-llm-eval Slovenian LLM Evaluation Dataset This dataset is designed for evaluating Slovenian language models and builds upon the work of gordicaleksa/slovenian-llm-eval-v0 which translated some of the popular English benchmarks into Slovenian by using Google Translate. We have further improved the quality of the Slovenian translations. The dataset contains the following benchmarks: ARC Challenge ARC Easy BoolQ GSM8K HellaSwag NQ Open OpenBookQA PIQA TriviaQA TruthfulQA Winogrande… See the full description on the dataset page: https://huggingface.co/datasets/cjvt/slovenian-llm-eval.tabular100K<n<1M0 likes179 downloads5mo agoHugging Face02MinaGabriel /llm-fol-reasoning-eval LLM FOL Reasoning Eval This dataset is derived from ProverQA, a First-Order Logic reasoning benchmark designed to test the ability of large language models (LLMs) to perform structured logical reasoning.It restructures and normalizes the ProverQA development and training data into a unified, clean format suitable for evaluating chain-of-thought (CoT) and symbolic reasoning capabilities in LLMs. Source Original dataset: ProverQA: A First-Order Logic Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/MinaGabriel/llm-fol-reasoning-eval.tabulartext-classification1K<n<10K3 likes120 downloads1y agoHugging Face03Spa-Bench /spa-bench-eval-rollouts-groot-n1-7-frozen-llm-full Spa-Bench Evaluation Rollouts — GR00T-N1.7 Frozen LLM (Full) Open this dataset in the LeRobot visualizer This repository contains physical Spa-Bench rollout trajectories with synchronized middle and wrist RGB video, robot state, action, timestamps, episode indices, and task indices. The native Hugging Face Data Studio viewer is enabled through the Parquet files declared above. Dataset summary Coverage: 900 scored recordings. Format: LeRobot v3 at 30 FPS. Task… See the full description on the dataset page: https://huggingface.co/datasets/Spa-Bench/spa-bench-eval-rollouts-groot-n1-7-frozen-llm-full.tabularrobotics100K<n<1M0 likes82 downloads12d agoHugging Face04nbvbharath-1729 /llm-math-evaluation-dataset LLM Math Response Evaluation Dataset Dataset Summary A human-annotated dataset of 150 AI-generated math responses evaluated across GPT-4o, Claude, and Gemini. Each response is scored on Correctness, Reasoning, and Clarity using a structured rubric, with written justification for every score. Supported Tasks LLM evaluation and benchmarking Math reasoning quality assessment Error type classification in AI responses Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nbvbharath-1729/llm-math-evaluation-dataset.tabularn<1K0 likes71 downloads5d agoHugging Face05HaimingW /ptb-llmevalmedtabularn<1K0 likes70 downloads3mo agoHugging Face06bermaneh /pde-llm-eval-code-perturbation-dataset pde-llm-eval-code-perturbation-dataset Code Perturbation Dataset (final). 256 PDE solver implementations: 64 base programs (32 physically valid, 32 with an injected bug that still runs to completion) expanded by 4 lexical perturbation conditions each -- original comments, comments removed, comments swapped in from a different implementation, and descriptive identifiers replaced by meaningless placeholders. The perturbations change the lexical surface only; executable behaviour… See the full description on the dataset page: https://huggingface.co/datasets/bermaneh/pde-llm-eval-code-perturbation-dataset.tabularn<1K0 likes66 downloads29d agoHugging Face07vector-institute /llm-eval-requeststabularn<1K0 likes52 downloads2y agoHugging Face08bay-calibration-llm-evaluators /summeval-annotated-latest SummEval-LLMEval Dataset Overview The original SummEval dataset (Fabbri et al., 2021) consists of 1,600 summaries annotated by human expert evaluators using a 5-point Likert scale across 4 criteria: coherence, consistency, fluency, and relevance. These 1,600 summaries are based on 100 source articles from the CNN/DailyMail dataset (Hermann et al., 2015). For each source article, SummEval collects 16 summaries generated by 16 different automatic summarization systems. Each… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/summeval-annotated-latest.tabular1K<n<10K0 likes48 downloads2y agoHugging Face09wayne-redemption /Sensor_Driven_Environmental_Monitoring_LLM_Evaluation_Dataset 📌 Dataset Contents Each sample includes: category: The evaluation domain prompt: The question given to the LLM temperature: Environmental temperature input humidity: Environmental humidity input context: A scenario label (e.g., cool_humid, hot_dry, average_day) reference: Expert-crafted expected output All data is provided in a single JSON file. 🧪 Intended Use This dataset supports research on: LLM evaluation methods (semantic similarity, contextual… See the full description on the dataset page: https://huggingface.co/datasets/wayne-redemption/Sensor_Driven_Environmental_Monitoring_LLM_Evaluation_Dataset.tabulartext-classificationn<1K0 likes48 downloads10mo agoHugging Face10AITrailblazer /repro-efficient-inference-for-noisy-llm-as-a-judge-evaluation-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K2 likes44 downloads2mo agoHugging Face11g-for-gour /llm-commit-message-evaluation Dataset Card for LLM Commit Message Evaluation The LLM Commit Message Evaluation dataset is designed to evaluate and compare the performance of Large Language Models (LLMs) in generating high-quality git commit messages. It contains real-world code diffs, issue descriptions, and issue titles extracted from open-source repositories (such as OWASP/Nest). For each code change, the dataset provides the original human-written commit message alongside commit messages generated by… See the full description on the dataset page: https://huggingface.co/datasets/g-for-gour/llm-commit-message-evaluation.tabularn<1K0 likes44 downloads2mo agoHugging Face12danielNisnevich /LLM-Eval-Filteredtabular1M<n<10M0 likes37 downloads2y agoHugging Face13amilmshaji /onepane-llm-evaluation-geminitabularn<1K0 likes36 downloads2y agoHugging Face14bay-calibration-llm-evaluators /mtbench-annotated-latest MT-Bench-Select Dataset Introduction The MT-Bench-Select dataset is a refined subset of the original MT-Bench dataset introduced by Zheng et al. (2023). The original MT-Bench dataset comprises 80 questions with answers generated by six models. Each question and each pair of models form an evaluation task, resulting in 1,200 tasks. For this dataset, we used a curated subset of the original MT-Bench dataset, as prepared by the authors of the LLMBar paper (Zeng et al.… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/mtbench-annotated-latest.tabular1K<n<10K0 likes32 downloads2y agoHugging Face15rubricreward /llm-metric-mm-eval-pairwisetabular1K<n<10K0 likes30 downloads1y agoHugging Face16llm-compe-2025-kato /step2-evaluated-dataset-test2 Complete Evaluation Dataset (Rubric + LogP) This dataset contains chain-of-thought explanations evaluated using both comprehensive rubric assessment and LogP evaluation. Overview Source Dataset: llm-compe-2025-kato/step2-evaluated-dataset-test2 Total Samples: 92 Successfully Evaluated (Rubric): 92 Failed Evaluations (Rubric): 0 Evaluation Model: Qwen/Qwen3-32B Rubric Evaluation Results Average Rubric Scores (0-4 scale) logical_coherence: 3.51… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-evaluated-dataset-test2.tabulartext-generationn<1K0 likes30 downloads1y agoHugging Face17danielNisnevich /LLM_Eval_Small_Exampletabular1M<n<10M0 likes29 downloads2y agoHugging Face18bay-calibration-llm-evaluators /llmbar-annotated-latest LLMBar-Select Dataset Introduction The LLMBar-Select dataset is a curated subset of the original LLMBar dataset introduced by Zeng et al. (2024). The LLMBar dataset consists of 419 instances, each containing an instruction paired with two outputs: one that faithfully follows the instruction and another that deviates while presenting superficially appealing qualities. It is designed to evaluate LLM-based evaluators more rigorously and objectively than previous benchmarks.… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/llmbar-annotated-latest.tabular1K<n<10K0 likes28 downloads2y agoHugging Face19Jake /dv-llm-eval-resultstabularn<1K0 likes28 downloads4mo agoHugging Face20open-llm-leaderboard /inumulaisk__eval_model-detailsgated Dataset Card for Evaluation run of inumulaisk/eval_model Dataset automatically created during the evaluation run of model inumulaisk/eval_model The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/inumulaisk__eval_model-details.tabular10K<n<100K0 likes27 downloads2y agoHugging Face21llm-compe-2025-kato /step2-evaluated-dataset-Qwen3-14B-cp32 Complete Evaluation Dataset (Rubric + LogP) This dataset contains chain-of-thought explanations evaluated using both comprehensive rubric assessment and LogP evaluation. Overview Source Dataset: llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B-cp32 Total Samples: 60 Successfully Evaluated (Rubric): 53 Failed Evaluations (Rubric): 7 Evaluation Model: Qwen/Qwen3-32B Rubric Evaluation Results Average Rubric Scores (0-4 scale) logical_coherence:… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B-cp32.tabulartext-generationn<1K0 likes25 downloads1y agoHugging Face22ku-nlp /jp-llm-evaluator-training Japanese LLM Evaluator Training Dataset It realeased on NLP2025 Constructing Open-source Large Language Model Evaluator for Japanese Overview Japanese LLM Evaluator Training Dataset is a dataset using for training Japanese LLM evaluator, which is focus on evaluate Japanese LLM from mutiple perspectives and meeting diverse evaluation requirements. Content The dataset includes 1000 diveser score rubrics. For every score rubrics, we generate 20 different… See the full description on the dataset page: https://huggingface.co/datasets/ku-nlp/jp-llm-evaluator-training.tabular10K<n<100K0 likes23 downloads2y agoHugging Face23llm-compe-2025-kato /step2-evaluated-dataset-Qwen3-14B-cp40 Complete Evaluation Dataset (Rubric + LogP) This dataset contains chain-of-thought explanations evaluated using both comprehensive rubric assessment and LogP evaluation. Overview Source Dataset: llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B-cp40 Total Samples: 58 Successfully Evaluated (Rubric): 53 Failed Evaluations (Rubric): 5 Evaluation Model: Qwen/Qwen3-32B Rubric Evaluation Results Average Rubric Scores (0-4 scale) logical_coherence:… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B-cp40.tabulartext-generationn<1K0 likes23 downloads1y agoHugging Face24sergiogpinto /memefact-llm-evaluations MemeFact LLM Evaluations Dataset This dataset contains 7,680 evaluation records where state-of-the-art Large Language Models (LLMs) assessed fact-checking memes according to specific quality criteria. The dataset provides comprehensive insights into how different AI models evaluate visual-textual content and how these evaluations compare to human judgments. Dataset Description Overview The "MemeFact LLM Evaluations" dataset documents a systematic… See the full description on the dataset page: https://huggingface.co/datasets/sergiogpinto/memefact-llm-evaluations.image1K<n<10K0 likes22 downloads1y agoHugging Face25cemig-ceia-v2 /energy_D_eval_llm_as_judge_granular_v6tabularn<1K0 likes21 downloads2mo agoHugging Face26llm-compe-2025-kato /step2-evaluated-dataset-Qwen3-14B Complete Evaluation Dataset (Rubric + LogP) This dataset contains chain-of-thought explanations evaluated using both comprehensive rubric assessment and LogP evaluation. Overview Source Dataset: llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B Total Samples: 156 Successfully Evaluated (Rubric): 135 Failed Evaluations (Rubric): 21 Evaluation Model: Qwen/Qwen3-32B Rubric Evaluation Results Average Rubric Scores (0-4 scale) logical_coherence:… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B.tabulartext-generationn<1K0 likes20 downloads1y agoHugging Face27llm-aes /dataset_hanna_96_prompts_llm_evaltabular1K<n<10K0 likes15 downloads3y agoHugging Face28llm-aes /llmeval2-annotated-latesttabular1K<n<10K0 likes15 downloads2y agoHugging Face29bay-calibration-llm-evaluators /llmeval2-annotated-latest LLMEval²-Select Dataset Introduction The LLMEval²-Select dataset is a curated subset of the original LLMEval² dataset introduced by Zhang et al. (2023). The original LLMEval² dataset comprises 2,553 question-answering instances, each annotated with human preferences. Each instance consists of a question paired with two answers. To construct LLMEval²-Select, Zeng et al. (2024) followed these steps: Labelled each instance with the human-preferred answer. Removed all… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/llmeval2-annotated-latest.tabular1K<n<10K0 likes15 downloads2y agoHugging Face30CoreyMorris /hugging-face-LLM-evaluation-results Dataset Summary Data comes from hugging face evaluation results using the https://github.com/EleutherAI/lm-evaluation-harness . See https://huggingface.co/datasets/open-llm-leaderboard/results for full results. tabularn<1K0 likes14 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.