CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nbvbharath-1729 /llm-math-evaluation-dataset LLM Math Response Evaluation Dataset Dataset Summary A human-annotated dataset of 150 AI-generated math responses evaluated across GPT-4o, Claude, and Gemini. Each response is scored on Correctness, Reasoning, and Clarity using a structured rubric, with written justification for every score. Supported Tasks LLM evaluation and benchmarking Math reasoning quality assessment Error type classification in AI responses Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nbvbharath-1729/llm-math-evaluation-dataset.tabularn<1K0 likes71 downloads5d agoHugging Face02g-for-gour /llm-commit-message-evaluation Dataset Card for LLM Commit Message Evaluation The LLM Commit Message Evaluation dataset is designed to evaluate and compare the performance of Large Language Models (LLMs) in generating high-quality git commit messages. It contains real-world code diffs, issue descriptions, and issue titles extracted from open-source repositories (such as OWASP/Nest). For each code change, the dataset provides the original human-written commit message alongside commit messages generated by… See the full description on the dataset page: https://huggingface.co/datasets/g-for-gour/llm-commit-message-evaluation.tabularn<1K0 likes44 downloads2mo agoHugging Face03AmirMohseni /LLM_Response_Eval LLM Response Evaluation Dataset This dataset contains a collection of responses generated by three large language models (LLMs): GPT-4o, Gemini 1.5 Pro, and Llama 3.1 405B. The responses are to a series of questions aimed at evaluating the models' problem-solving abilities using Polya's problem-solving technique, as described in the book "How to Solve It" by George Polya. Dataset Overview Questions: The dataset includes a set of questions designed to evaluate the… See the full description on the dataset page: https://huggingface.co/datasets/AmirMohseni/LLM_Response_Eval.textquestion-answeringn<1K0 likes27 downloads2y agoHugging Face04sergiogpinto /memefact-llm-evaluations MemeFact LLM Evaluations Dataset This dataset contains 7,680 evaluation records where state-of-the-art Large Language Models (LLMs) assessed fact-checking memes according to specific quality criteria. The dataset provides comprehensive insights into how different AI models evaluate visual-textual content and how these evaluations compare to human judgments. Dataset Description Overview The "MemeFact LLM Evaluations" dataset documents a systematic… See the full description on the dataset page: https://huggingface.co/datasets/sergiogpinto/memefact-llm-evaluations.image1K<n<10K0 likes22 downloads1y agoHugging Face05CoreyMorris /hugging-face-LLM-evaluation-results Dataset Summary Data comes from hugging face evaluation results using the https://github.com/EleutherAI/lm-evaluation-harness . See https://huggingface.co/datasets/open-llm-leaderboard/results for full results. tabularn<1K0 likes14 downloads3y agoHugging Face06userhuggingface4321 /multilingual-llm-evaluation Multilingual LLM Evaluation A small evaluation dataset for comparing language models across English, Hindi, and Spanish. Columns language: language code (en, hi, or es) question: question provided to the model expected_answer: reference answer used for scoring Intended use This dataset can be used to compare model accuracy, language adherence, and response speed across languages. Limitations This is a small demonstration dataset and… See the full description on the dataset page: https://huggingface.co/datasets/userhuggingface4321/multilingual-llm-evaluation.textquestion-answeringn<1K0 likes12 downloads2mo agoHugging Face07strikoder /LLM-EvaluationHub LLM-EvaluationHub: Enhanced Dataset for Large Language Model Assessment This repository, LLM-EvaluationHub, presents an enhanced dataset tailored for the evaluation and assessment of Large Language Models (LLMs). It builds upon the dataset originally provided by SafetyBench (THU-COAI), incorporating significant modifications and additions to address specific research objectives. Below is a summary of the key differences and enhancements: Key Modifications… See the full description on the dataset page: https://huggingface.co/datasets/strikoder/LLM-EvaluationHub.textzero-shot-classification1K<n<10K1 likes11 downloads3y agoHugging Face08sghosts /ar-llm-vibe-evaltextn<1K0 likes10 downloads1y agoHugging Face09SMARTICT /Pubmed-RAG-TR-LLM-EvalLLM-as-a-judge evalaution results using "claude-haiku-4-5-20251001" for SMARTICT/Pubmed-RAG-TR-LLM dataset. tabular1K<n<10K0 likes8 downloads7mo agoHugging Face10hareem-arshad /llm-reasoning-evaluation-examples LLM Reasoning Evaluation Examples Overview This dataset contains 20 prompts designed to evaluate potential blind spots in a base language model.It covers 10 reasoning categories: Logic, Arithmetic, Factual, Language/Translation, Pattern/Analogy, Commonsense, Comparison, Negation/Uncertainty, Math Sequence, and Analogies. The dataset is intended for educational and research purposes to illustrate where a base model may produce outputs that differ from expected answers.… See the full description on the dataset page: https://huggingface.co/datasets/hareem-arshad/llm-reasoning-evaluation-examples.textn<1K0 likes8 downloads7mo agoHugging Face11davanstrien /llm-pubmed-query-generation-evaltabularn<1K0 likes6 downloads2y agoHugging Face12MahdiAbdoZahra /LLM_EVAL_Datasetgatedtext10K<n<100K0 likes5 downloads2y agoHugging Face13MetalZuna /simple_llm_eval_frameworktextn<1K0 likes4 downloads3y agoHugging Face14renataaraujoe /Bilingual-LLM-Eval-106 Bilingual-LLM-Eval-106 📌 Overview Bilingual-LLM-Eval-106 is a curated, human-annotated evaluation dataset of 106 LLM response pairs in English and Portuguese, designed to benchmark model performance across multiple quality dimensions. The dataset focuses on realistic and challenging evaluation scenarios, including adversarial prompts, ambiguous queries, and hard negatives. It is intended for: LLM evaluation and benchmarking LLM-as-a-Judge research Multilingual… See the full description on the dataset page: https://huggingface.co/datasets/renataaraujoe/Bilingual-LLM-Eval-106.tabularn<1K0 likes4 downloads6mo agoHugging Face15Pravin908 /llm_evaluationtabularn<1K0 likes2 downloads1y agoHugging Face16Kasaf /llm-eval-recipe-impacttextn<1K0 likes1 downloads2y agoHugging Face17evapaunova /toy-llm-eval-datasettextn<1K0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.