CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nbvbharath-1729 /llm-math-evaluation-dataset LLM Math Response Evaluation Dataset Dataset Summary A human-annotated dataset of 150 AI-generated math responses evaluated across GPT-4o, Claude, and Gemini. Each response is scored on Correctness, Reasoning, and Clarity using a structured rubric, with written justification for every score. Supported Tasks LLM evaluation and benchmarking Math reasoning quality assessment Error type classification in AI responses Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nbvbharath-1729/llm-math-evaluation-dataset.tabularn<1K0 likes62 downloads3d agoHugging Face02g-for-gour /llm-commit-message-evaluation Dataset Card for LLM Commit Message Evaluation The LLM Commit Message Evaluation dataset is designed to evaluate and compare the performance of Large Language Models (LLMs) in generating high-quality git commit messages. It contains real-world code diffs, issue descriptions, and issue titles extracted from open-source repositories (such as OWASP/Nest). For each code change, the dataset provides the original human-written commit message alongside commit messages generated by… See the full description on the dataset page: https://huggingface.co/datasets/g-for-gour/llm-commit-message-evaluation.tabularn<1K0 likes51 downloads2mo agoHugging Face03wayne-redemption /Sensor_Driven_Environmental_Monitoring_LLM_Evaluation_Dataset 📌 Dataset Contents Each sample includes: category: The evaluation domain prompt: The question given to the LLM temperature: Environmental temperature input humidity: Environmental humidity input context: A scenario label (e.g., cool_humid, hot_dry, average_day) reference: Expert-crafted expected output All data is provided in a single JSON file. 🧪 Intended Use This dataset supports research on: LLM evaluation methods (semantic similarity, contextual… See the full description on the dataset page: https://huggingface.co/datasets/wayne-redemption/Sensor_Driven_Environmental_Monitoring_LLM_Evaluation_Dataset.tabulartext-classificationn<1K0 likes50 downloads10mo agoHugging Face04AITrailblazer /repro-efficient-inference-for-noisy-llm-as-a-judge-evaluation-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K2 likes40 downloads2mo agoHugging Face05amilmshaji /onepane-llm-evaluation-geminitabularn<1K0 likes34 downloads2y agoHugging Face06sergiogpinto /memefact-llm-evaluations MemeFact LLM Evaluations Dataset This dataset contains 7,680 evaluation records where state-of-the-art Large Language Models (LLMs) assessed fact-checking memes according to specific quality criteria. The dataset provides comprehensive insights into how different AI models evaluate visual-textual content and how these evaluations compare to human judgments. Dataset Description Overview The "MemeFact LLM Evaluations" dataset documents a systematic… See the full description on the dataset page: https://huggingface.co/datasets/sergiogpinto/memefact-llm-evaluations.image1K<n<10K0 likes23 downloads1y agoHugging Face07CoreyMorris /hugging-face-LLM-evaluation-results Dataset Summary Data comes from hugging face evaluation results using the https://github.com/EleutherAI/lm-evaluation-harness . See https://huggingface.co/datasets/open-llm-leaderboard/results for full results. tabularn<1K0 likes15 downloads3y agoHugging Face08amilmshaji /onepane-llm-evaluationtabularn<1K0 likes12 downloads2y agoHugging Face09amilmshaji /llm-evaluationtabularn<1K0 likes7 downloads2y agoHugging Face10Pravin908 /llm_evaluationtabularn<1K0 likes2 downloads1y agoHugging Face11eliyahabba /llm-evaluation-analysisgatedtabular100M<n<1B0 likes1 downloads2y agoHugging Face12eliyahabba /llm-evaluation-analysis-splitgatedtabular100M<n<1B0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.