CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CohereLabs /aya_evaluation_suite Dataset Summary Aya Evaluation Suite contains a total of 26,750 open-ended conversation-style prompts to evaluate multilingual open-ended generation quality.To strike a balance between language coverage and the quality that comes with human curation, we create an evaluation suite that includes: human-curated examples in 7 languages (tur, eng, yor, arb, zho, por, tel) → aya-human-annotated. machine-translations of handpicked examples into 101 languages → dolly-machine-translated.… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_evaluation_suite.tabulartext-generation10K<n<100K55 likes6.7k downloads1y agoHugging Face02zjunlp /Chat2Workflow-Evaluation Chat2Workflow Chat2Workflow is a benchmark designed for evaluating the ability of Large Language Models (LLMs) to generate executable visual workflows from natural language instructions. Paper: Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language Repository: zjunlp/Chat2Workflow Overview Executable visual workflows are widely used in industrial deployments for their reliability and controllability. Chat2Workflow addresses the… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/Chat2Workflow-Evaluation.documenttext-generationn<1K4 likes3.1k downloads4mo agoHugging Face03aisingapore /Cultural-Evaluation-Kalahigated Kalahi Kalahi evaluates the ability of LLMs to generate responses relevant to Filipino culture in terms of shared knowledge and ethics. This dataset contains a MCQ-compatible version of the Kalahi dataset that is used in SEA-HELM. Supported Tasks and Leaderboards Kalahi is designed for evaluating Filipino cultural representations in instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore. Languages Tagalog (tl)… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Cultural-Evaluation-Kalahi.textmultiple-choicen<1K0 likes890 downloads9mo agoHugging Face04tsinghua-sigs-robot-lab /VeriLoop-E2-Evaluation-Evidence VeriLoop E2 Evaluation Evidence Public evaluation evidence for VeriLoop E2 across nine code, agentic, mathematical, and scientific reasoning benchmarks. This dataset repository is the canonical public evidence layer for the reported benchmark results of VeriLoop E2, a post-trained model based on Qwen 3.8-27B. It is designed to separate headline benchmark reporting from the underlying auditable artifacts required to inspect, reproduce, and verify those results. The repository… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence.texttext-generation0 likes818 downloads2d agoHugging Face05JesseLiu /patient-evaluations Patient Evaluations Dataset This dataset contains clinician evaluations of AI-generated patient summaries from MIMIC-III data. Dataset Description The dataset includes expert clinician assessments of AI-generated patient summaries, with detailed ratings across multiple dimensions including clinical accuracy, completeness, relevance, and identification of hallucinations or critical omissions. Dataset Structure The dataset contains a CSV file… See the full description on the dataset page: https://huggingface.co/datasets/JesseLiu/patient-evaluations.texttext-generationn<1K0 likes661 downloads7mo agoHugging Face06YixuanEvenXu /HIP-training-and-evaluation-data HIP Training and Evaluation Data This dataset contains the text data released with Base Models Look Human To AI Detectors for reproducing the Humanization by Iterative Paraphrasing (HIP) training setup and the prefix-based continuation evaluation. Configs training data/train.parquet contains 10,581 supervised HIP training pairs with seven columns: dataset: upstream dataset family, either raid or mage. source: selected source domain or subcorpus. text: original… See the full description on the dataset page: https://huggingface.co/datasets/YixuanEvenXu/HIP-training-and-evaluation-data.tabulartext-generation10K<n<100K1 likes350 downloads4mo agoHugging Face07MultiverseComputingCAI /llm-refusal-evaluation 🛡️ LLM Refusal Evaluation Benchmark This repository contains the benchmarks used in the LLM-Refusal-Evaluation suite. The prompts are organized into three groups: Safety Benchmarks — harmful / jailbreak-style prompts that models should refuse. Chinese Sensitive Topics — prompts that may be censored by China-aligned models. Sanity Check Datasets — non-sensitive prompts to ensure models don’t over-refuse. 📌 Contents Safety Benchmarks JailbreakBench SorryBench… See the full description on the dataset page: https://huggingface.co/datasets/MultiverseComputingCAI/llm-refusal-evaluation.texttext-generation1K<n<10K5 likes244 downloads9mo agoHugging Face08compass-group-tue /sdf_evaluation_traits Models That Know How Evaluations Are Designed Score Safer This repository contains the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer. Project Page | GitHub Repository Dataset Description These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations. Documents were generated using the… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits.tabulartext-generation10K<n<100K1 likes201 downloads4mo agoHugging Face09openbmb /RLPR-Evaluation Dataset Card for RLPR-Evaluation GitHub | Paper News: [2025.06.23] 📃 Our paper detailing the RLPR framework and its comprehensive evaluation using this suite is accessible at here! Dataset Summary We include the following seven benchmarks for evaluation of RLPR: Mathematical Reasoning Benchmarks: MATH-500 (Cobbe et al., 2021) Minerva (Lewkowycz et al., 2022) AIME24 General Domain Reasoning Benchmarks: MMLU-Pro (Wang et al., 2024): A multitask language… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/RLPR-Evaluation.textquestion-answeringn<1K3 likes196 downloads1y agoHugging Face10TsinghuaC3I /ZEDA-Evaluation ZEDA Dataset This repository contains the training data for ZEDA (Zero-Expert Self-Distillation Adaptation), a framework introduced in the paper Post-Trained MoE Can Skip Half Experts via Self-Distillation. ZEDA is a low-cost framework that transforms post-trained static MoE models into efficient dynamic ones by injecting zero experts and using self-distillation. Paper: Post-Trained MoE Can Skip Half Experts via Self-Distillation GitHub Repository:… See the full description on the dataset page: https://huggingface.co/datasets/TsinghuaC3I/ZEDA-Evaluation.texttext-generation0 likes185 downloads4mo agoHugging Face11HeAAAAA /Crab-role-playing-evaluation-benchmark 📄 Paper | 📄 Github 💬 Role-playing Model | 💬 Role-palying Evaluation Model 💬 Training Dataset | 💬 Evaluation Benchmark | 💬 Annotated Role-playing Evaluation Dataset | 💬 Human-preference Dataset 1. Introduction This is the dataset used for evalauating a role‑playing LLM. More details can be seen at GitHub and Crab… See the full description on the dataset page: https://huggingface.co/datasets/HeAAAAA/Crab-role-playing-evaluation-benchmark.texttext-generationn<1K0 likes129 downloads1y agoHugging Face12jizej /Competence-Based-Evaluation Competence-Based Evaluation (Invariance Benchmark) A benchmark for testing whether language models give the same answer to semantically equivalent reformulations of a logical-ordering question. Given a set of pairwise constraints (e.g. Alice is in front of Bob), a model should answer transitive-closure queries (Is Carol in front of Dave?) consistently whether the constraints are stated using a relation or its inverse. Each item exists as a paired (original, equivalent) record… See the full description on the dataset page: https://huggingface.co/datasets/jizej/Competence-Based-Evaluation.textquestion-answering100K<n<1M0 likes126 downloads5mo agoHugging Face13Lots-of-LoRAs /task1338_peixian_equity_evaluation_corpus_sentiment_classifier Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1338_peixian_equity_evaluation_corpus_sentiment_classifier Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1338_peixian_equity_evaluation_corpus_sentiment_classifier.texttext-generation1K<n<10K0 likes109 downloads2y agoHugging Face14furkankarli /turkish-brand-bias-evaluations Turkish Brand Bias Evaluations / Türkçe Marka Yanlılığı Değerlendirmeleri Furkan Karlı tarafından Türkçe ürün ve hizmet önerilerindeki marka görünürlüğünü incelemek amacıyla oluşturulmuş LLM değerlendirme veri setidir. An LLM evaluation dataset curated by Furkan Karlı to study brand visibility in Turkish product and service recommendations. Veri seti özeti 300 tamamlanmış ve judge edilmiş yanıt Domainler: VPN (150) ve kozmetik (150) Koşullar: web araması kapalı… See the full description on the dataset page: https://huggingface.co/datasets/furkankarli/turkish-brand-bias-evaluations.tabulartext-generationn<1K1 likes101 downloads20d agoHugging Face15dougdotcon /douvras-ptbr-enterprise-ai-evaluation Douvras PT-BR Enterprise AI Evaluation Benchmark pequeno e auditável para avaliar assistentes empresariais em português brasileiro. A versão 0.2 adiciona um corpus de treino sintético separado; os splits de avaliação continuam congelados. O benchmark mede quatro capacidades que aparecem em projetos reais de IA aplicada: responder somente a partir de um documento fornecido; reconhecer quando a informação não está disponível; resistir a instruções maliciosas inseridas no conteúdo… See the full description on the dataset page: https://huggingface.co/datasets/dougdotcon/douvras-ptbr-enterprise-ai-evaluation.textquestion-answering0 likes97 downloads12d agoHugging Face16BushNate /llm-refusal-evaluation 🛡️ LLM Refusal Evaluation Benchmark This repository contains the benchmarks used in the LLM-Refusal-Evaluation suite. The prompts are organized into three groups: Safety Benchmarks — harmful / jailbreak-style prompts that models should refuse. Chinese Sensitive Topics — prompts that may be censored by China-aligned models. Sanity Check Datasets — non-sensitive prompts to ensure models don’t over-refuse. 📌 Contents Safety Benchmarks JailbreakBench SorryBench… See the full description on the dataset page: https://huggingface.co/datasets/BushNate/llm-refusal-evaluation.texttext-generation1K<n<10K1 likes88 downloads2mo agoHugging Face17316usman /llm-output-evaluation LLM_OUTPUT_EVALUATION A preference dataset for LLM_OUTPUT_EVALUATION, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate. Format Standard preference / DPO schema — each row: column meaning prompt the request (originally prompt) chosen the human-preferred response rejected a worse response to the same prompt source the dataset/URL the row was harvested from Splits… See the full description on the dataset page: https://huggingface.co/datasets/316usman/llm-output-evaluation.texttext-generation1K<n<10K0 likes69 downloads15d agoHugging Face18WueNLP /mHallucination_Evaluation Multilingual Hallucination Evaluation in the wild The dataset was as part of the paper: How Much Do LLMs Hallucinate across Languages? On Multilingual Estimation of LLM Hallucination in the Wild Below is the figure summarizing the multilingual hallucination evaluation dataset creation (and multilingual hallucination detection dataset): Dataset Details The dataset is a high quality synthetic query/prompt and wikipedia reference pair for estimating hallucinations in the… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/mHallucination_Evaluation.texttext-generation10K<n<100K0 likes64 downloads2y agoHugging Face19aiagentkarl /agent-evaluation-benchmark Agent Evaluation Benchmark A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories. Overview This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more. Categories Category Test Cases Description Weather 5 Forecasts, UV index, climate history Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.texttext-generationn<1K0 likes58 downloads6mo agoHugging Face20Chemin-AI /advent_of_code_evaluations Advent of Code Evaluation This evaluation is conducted on the advent of code dataset on several models including Qwen2.5-Coder-32B-Instruct, DeepSeek-V3-fp8, Llama-3.3-70B-Instruct, GPT-4o-mini, DeepSeek-R1.The aim is to to see how well these models can handle real-world puzzle prompts, generate correct Python code, and ultimately shed light on which LLM truly excels at reasoning and problem-solving.We used pass@1 to measure the functional correctness. Results… See the full description on the dataset page: https://huggingface.co/datasets/Chemin-AI/advent_of_code_evaluations.texttext-generationn<1K2 likes50 downloads2y agoHugging Face21iammytoo /japanese-humor-evaluation-v2 Japanese Multimodal Humor Evaluation Dataset (v2) 画像/テキストのお題に対する面白い回答のデータセット。bokete(画像→テキスト)とkeitai(テキスト→テキスト)を統合。 使い方 from datasets import load_dataset dataset = load_dataset("iammytoo/japanese-humor-evaluation-v2") データ構造 odai_type: 'image' or 'text' image: 画像お題(textタイプではNone) odai: テキストお題(imageタイプではNone) response: 回答テキスト score: 0-4の正規化スコア ソース YANS-official/ogiri-bokete YANS-official/ogiri-keitai imagetext-generation10K<n<100K0 likes50 downloads1y agoHugging Face22AnjanSB /NQ-RAG-DPO-Evaluation Dataset Card Dataset Summary This repository contains five structured data files forming a complete evaluation and training workflow for a multi-perspective alignment system built on Retrieval-Augmented Generation (RAG) and Direct Preference Optimization (DPO). The system is organized into three interconnected pipelines: 1️. RAG Pipeline The RAG pipeline uses the Natural Questions (validation split) as the knowledge source and evaluation benchmark. For each… See the full description on the dataset page: https://huggingface.co/datasets/AnjanSB/NQ-RAG-DPO-Evaluation.texttext-generation1K<n<10K1 likes50 downloads7mo agoHugging Face23Experimental-Orange /HumanAgencyBench_Evaluation_Results HumanAgencyBench evaluation results Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants Code: https://github.com/BenSturgeon/HumanAgencyBench/ Dataset Description This dataset contains comprehensive evaluation results from testing 25 different language models across 6 areas of behaviours critical for human agency support. Each model was evaluated on 3,000 prompts (500 per category), resulting in 75,000 total evaluations designed… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Evaluation_Results.tabulartext-generation10K<n<100K0 likes47 downloads26d agoHugging Face24SahmBenchmark /fatwa-qa-evaluation Fatwa QA Evaluation Dataset Dataset Description This dataset contains Islamic finance and jurisprudence fatwa question-answer pairs for evaluating Arabic language models. This is an open-ended QA evaluation benchmark where models generate free-form answers. Dataset Statistics Total Samples: 2,000 Average Question Length: 243.9 characters Average Answer Length: 492.3 characters Dataset Structure Data Fields id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/fatwa-qa-evaluation.tabularquestion-answering1K<n<10K0 likes44 downloads10mo agoHugging Face25compass-group-tue /sdf_evaluation_traits_15M Models That Know How Evaluations Are Designed Score Safer This repository contains a subset of the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer. Project Page | GitHub Repository Dataset Description These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations. Documents were generated… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits_15M.tabulartext-generation10K<n<100K0 likes35 downloads1mo agoHugging Face26jayzou3773 /less-is-moe-gpqa-diamond-evaluation Less-is-MoE GPQA-Diamond evaluation set This private dataset stores the 198-question GPQA-Diamond evaluation file used by the MoE-Honing evaluation format. Upstream source: Idavidrein/gpqa, config gpqa_diamond Upstream revision: 633f5ee89ab8ad4522a9f850766b73f62147ffdd Split: test Rows: 198 SHA-256: d5b0d6dad6c1993a8cb17fd7aa635fbc5e8b684ae42b5b6e80d15467b399eec3 Fields: problem, solution, domain The problem field contains the formatted four-choice prompt, solution stores the… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-gpqa-diamond-evaluation.textquestion-answeringn<1K0 likes35 downloads5d agoHugging Face27yunjaeys /Contextual_Response_Evaluation_for_ESL_and_ASD_Support Dataset Card for "Contextual Response Evaluation for ESL and ASD Support💜💬🌐"" Dataset Description 📖 Dataset Summary 📝 Curated by Eric Soderquist, this dataset is a collection of English prompts and responses generated by the Phi-2 model, designed to evaluate and improve NLP models for supporting ESL (English as a Second Language) and ASD (Autism Spectrum Disorder) user bases. Each prompt is paired with multiple AI-generated responses and evaluated using a… See the full description on the dataset page: https://huggingface.co/datasets/yunjaeys/Contextual_Response_Evaluation_for_ESL_and_ASD_Support.texttext-generationn<1K0 likes26 downloads3y agoHugging Face28taln-ls2n /keyphrase_homogeneity_evaluation license: cc-by-nc-4.0 language: - en size_categories: - n<1K Data pairs used in the evaluation of the paper "[Evaluating the Homogeneity of Keyphrase Prediction Models]"(https://arxiv.org/abs/2602.12989), Maël Houbre, Florian Boudin and Béatrice Daille, LREC 2026 texttext-generation10K<n<100K0 likes26 downloads7mo agoHugging Face29INTERX /Molding-Generation-Evaluation Molding-Generation-Evaluation Dataset Molding 도메인(사출 성형, 금형)에 대한 LLM의 생성 품질을 평가하기 위한 데이터셋입니다. texttext-generationn<1K0 likes24 downloads1y agoHugging Face30hadeelbkh /tokenized-IELTS-writing-task-2-evaluation-DialoGPT-mediumtexttext-generation1K<n<10K2 likes23 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.