CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MERA-evaluation /MERA MERA (Multimodal Evaluation for Russian-language Architectures) Summary MERA (Multimodal Evaluation for Russian-language Architectures) is a new open independent benchmark for the evaluation of SOTA models for the Russian language. The MERA benchmark unites industry and academic partners in one place to research the capabilities of fundamental models, draw attention to AI-related issues, foster collaboration within the Russian Federation and in the international arena… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/MERA.text10K<n<100K11 likes3.1k downloads2y agoHugging Face02zjunlp /Chat2Workflow-Evaluation Chat2Workflow Chat2Workflow is a benchmark designed for evaluating the ability of Large Language Models (LLMs) to generate executable visual workflows from natural language instructions. Paper: Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language Repository: zjunlp/Chat2Workflow Overview Executable visual workflows are widely used in industrial deployments for their reliability and controllability. Chat2Workflow addresses the… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/Chat2Workflow-Evaluation.documenttext-generationn<1K4 likes3.1k downloads4mo agoHugging Face03mmathys /openai-moderation-api-evaluation Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection" The evaluation dataset data/samples-1680.jsonl.gz is the test set used in this paper. Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label. Category Label Definition sexual S Content meant to arouse sexual… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/openai-moderation-api-evaluation.tabulartext-classification1K<n<10K38 likes2.9k downloads3y agoHugging Face04PKU-Alignment /BeaverTails-Evaluation Dataset Card for BeaverTails-Evaluation BeaverTails is an AI safety-focused collection comprising a series of datasets. This repository contains test prompts specifically designed for evaluating language model safety. It is important to note that although each prompt can be connected to multiple categories, only one category is labeled for each prompt. The 14 harm categories are defined as follows: Animal Abuse: This involves any form of cruelty or harm inflicted on animals… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/BeaverTails-Evaluation.texttext-classificationn<1K15 likes642 downloads3y agoHugging Face05minhhien0811 /java_evaluation_benchmarkstext1K<n<10K0 likes603 downloads4mo agoHugging Face06WenyiWU0111 /webvoyager_evaluation_datatextn<1K0 likes419 downloads1y agoHugging Face07FreedomIntelligence /Medical_Multimodal_Evaluation_Data Evaluation Guide This dataset is used to evaluate medical multimodal LLMs, as used in HuatuoGPT-Vision. It includes benchmarks such as VQA-RAD, SLAKE, PathVQA, PMC-VQA, OmniMedVQA, and MMMU-Medical-Tracks. To get started: Download the dataset and extract the images.zip file. Find evaluation code on our GitHub: HuatuoGPT-Vision. This open-source release aims to simplify the evaluation of medical multimodal capabilities in large models. Please cite the relevant benchmark… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical_Multimodal_Evaluation_Data.imageimage-to-text10K<n<100K29 likes346 downloads2y agoHugging Face08ZeeshanSaud /CodeTruthAgent-V3-Module1-Evaluation Source Code https://github.com/Zeeshan78699/CodeTruthAgent — tag v3.0.0-module1 CodeTruth Agent V3 — Module 1 Evaluation Validation results for Module 1: Repository Cognition Engine — a deterministic, rule-based engine that scans a software repository and determines its application type, primary framework, technology stack, and file inventory. What's in this dataset FULL_DOMAIN_SUMMARY.md — summary table of all 69 validated repositories… See the full description on the dataset page: https://huggingface.co/datasets/ZeeshanSaud/CodeTruthAgent-V3-Module1-Evaluation.textn<1K0 likes248 downloads3mo agoHugging Face09demisama /UGround-Offline-Evaluationimage1K<n<10K1 likes238 downloads2y agoHugging Face10compass-group-tue /sdf_evaluation_traits Models That Know How Evaluations Are Designed Score Safer This repository contains the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer. Project Page | GitHub Repository Dataset Description These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations. Documents were generated using the… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits.tabulartext-generation10K<n<100K1 likes204 downloads4mo agoHugging Face11orhunc /Bias-Evaluation-TurkishTranslation of bias evaluation framework of May et al. (2019) from this repository and this paper into Turkish. There is a total of 37 tests including tests addressing gender-bias as well as tests designed to evaluate the ethnic bias toward Kurdish people in Türkiye context. Abstract of the paper: While the growing size of pre-trained language models has led to large improvements in a variety of natural language processing tasks, the success of these models comes with a price: They are trained… See the full description on the dataset page: https://huggingface.co/datasets/orhunc/Bias-Evaluation-Turkish.textn<1K1 likes188 downloads4y agoHugging Face12gauravshrm211 /VC-startup-evaluation-for-investmentThis data set includes the completion pairs for evaluating startups before investing in them. This data set iincludes completion examples for Chain of Thought reasoning to perform financial calculations. This data set includes completion examples for evaluating risk profile, growth propspects, cost, ratios, market size, asset, liability, debt, equity and other ratios. This data set includes comparison of different startups. textn<1K13 likes160 downloads3y agoHugging Face13ltg /normistral-11b-thinking-evaluationtext10K<n<100K1 likes138 downloads10mo agoHugging Face14phionyx /airep-embedded-evaluation-profile AIREP Embedded Evaluation Profile v0.1 This is not a training dataset or benchmark. It is a Hugging Face distribution mirror of an experimental evaluation-evidence profile, its schema basis, registry and fixtures. The canonical specification history lives in the AIREP GitHub repository. Byte identity between this mirror and its source commit is a distribution-integrity property, not independent scientific verification. Experimental evaluation-evidence contract for AIREP v0.2.… See the full description on the dataset page: https://huggingface.co/datasets/phionyx/airep-embedded-evaluation-profile.textn<1K0 likes129 downloads5d agoHugging Face15jizej /Competence-Based-Evaluation Competence-Based Evaluation (Invariance Benchmark) A benchmark for testing whether language models give the same answer to semantically equivalent reformulations of a logical-ordering question. Given a set of pairwise constraints (e.g. Alice is in front of Bob), a model should answer transitive-closure queries (Is Carol in front of Dave?) consistently whether the constraints are stated using a relation or its inverse. Each item exists as a paired (original, equivalent) record… See the full description on the dataset page: https://huggingface.co/datasets/jizej/Competence-Based-Evaluation.textquestion-answering100K<n<1M0 likes128 downloads5mo agoHugging Face16ZurichNLP /romansh-mt-evaluation Dataset Description This dataset contains the results of a human evaluation of machine translations from German into the six Romansh varieties. The evaluations were carried out by native speakers of the respective Romansh idioms as well as professional linguists. The evaluation covers three quality dimensions: Document accuracy, in which annotators assessed the adequacy of complete document translations. Segment accuracy, in which annotators selected the more accurate… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/romansh-mt-evaluation.tabular1K<n<10K0 likes110 downloads2mo agoHugging Face17CaiZhiTech /Evaluation-Dataset-of-AI-Agent-Security-Guardrails DKnownAI Agent Security Evaluation Dataset Data Fields Field Type Description text string The adversarial input (prompt) to be evaluated by a security guardrail action string Human-annotated label: blocked or allowed Citation @misc{li2026comparativeevaluationaiagent, title={A Comparative Evaluation of AI Agent Security Guardrails}, author={Qi Li and Jiu Li and Pingtao Wei and Jianjun Xu and Xueyi Wei and Jiwei Shi and Xuan… See the full description on the dataset page: https://huggingface.co/datasets/CaiZhiTech/Evaluation-Dataset-of-AI-Agent-Security-Guardrails.texttext-classification1K<n<10K1 likes98 downloads5mo agoHugging Face18bingbangboom /stockfish-evaluation-SAN Dataset Card for the Stockfish Evaluations A dataset of chess positions evaluated with various flavours of Stockfish running within user browsers. Produced by, and for, the Lichess analysis board. Evaluations are formatted as JSON; one position per line. The schema of a position looks like this: { "fen": "8/8/2B2k2/p4p2/5P1p/Pb6/1P3KP1/8 w - -", "depth": 42, "evaluation": 5.64, "best_move": "Kg1", "best_line": "Kg1 Ke6 Kh2 Kd6 Be8 Kc5 Kh3 Kd6 Bb5 Ke7" } fen: string, the… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/stockfish-evaluation-SAN.text10M<n<100M3 likes81 downloads2y agoHugging Face19anikethh /Comparative-Idea-Evaluation 🔬 Comparative Idea Evaluation Dataset accompanying our Findings of ACL 2026 paper: Teaching Language Models to Forecast Research Success Through Comparative Idea Evaluation. Srujan P Mule · Aniketh Garikaparthi · Manasi Patwardhan 📄 ACL Anthology · arXiv Research overview from Figure 1 of the paper. This release contains the comparison datasets; reasoning-training variants illustrated in the figure are not included. 🧠 What is this dataset for? Given a research… See the full description on the dataset page: https://huggingface.co/datasets/anikethh/Comparative-Idea-Evaluation.texttext-classification10K<n<100K0 likes71 downloads7d agoHugging Face20SeanWang0027 /data_for_evaluationtabular1K<n<10K0 likes70 downloads6mo agoHugging Face21RKB109 /rag-evaluation-lab-20260829-dataset RAG Evaluation Lab Synthetic Dataset Summary This dataset contains 14 training examples and 4 held-out examples for RAG systems often ship without a stable regression set or failure taxonomy. Every record is synthetic and includes: input: query, event, or feature description label: expected class, route, relation, or evidence category context: synthetic supporting context source: fictional source identifier variant: generation pattern synthetic: always true… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/rag-evaluation-lab-20260829-dataset.texttext-classificationn<1K0 likes69 downloads25d agoHugging Face22llmsql-bench /benchmark-evaluation-resultstext1M<n<10M0 likes66 downloads7mo agoHugging Face23sabaridsnfuji /repro-accurate-evaluation-of-quickest-changepoint-detectors-via-non-parametric-survival-analysis Accurate Evaluation of Quickest Changepoint Detectors via Non-parametric Survival Analysis This is a reproduction logbook for ICML 2026. OpenReview ID: LhGxRnGmGJ Paper Abstract This logbook reproduces KM-ARL and KM-ADD estimators for changepoint detection. See logbook.json for full claim verification details. textn<1K0 likes66 downloads2mo agoHugging Face24Noothi /telugu-indicf5-evaluationaudion<1K0 likes65 downloads1mo agoHugging Face25s-emanuilov /rivers-evaluation-results Rivers Evaluation Results - Comprehensive LLM Benchmarking All results from the paper's five experimental conditions: baseline LLMs, fine-tuned models, RAG, and Graph-RAG with Licensing Oracle. This repository contains baseline evaluations for Claude Sonnet 4.5, Gemini 2.5 Flash Lite, and Gemma 3-4B, along with fine-tuning results for both factual recall and abstention behavior. It also includes outputs from the embedding-based RAG system and the Graph-RAG with Licensing Oracle… See the full description on the dataset page: https://huggingface.co/datasets/s-emanuilov/rivers-evaluation-results.text10K<n<100K0 likes58 downloads11mo agoHugging Face26RKB109 /rag-evaluation-lab-20260908-dataset RAG Evaluation Lab Synthetic Dataset Summary This dataset contains 14 training examples and 4 held-out examples for RAG systems often ship without a stable regression set or failure taxonomy. Every record is synthetic and includes: input: query, event, or feature description label: expected class, route, relation, or evidence category context: synthetic supporting context source: fictional source identifier variant: generation pattern synthetic: always true… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/rag-evaluation-lab-20260908-dataset.texttext-classificationn<1K0 likes58 downloads15d agoHugging Face27Reza-Telus /certainty-robustness-llm-evaluation Certainty Robustness Benchmark This repository accompanies the paper: Certainty robustness: Evaluating LLM stability under self-challenging promptsMohammadreza Saadat, Steve NemzerarXiv:2603.03330, 2026https://arxiv.org/abs/2603.03330 Overview The Certainty Robustness Benchmark evaluates how large language models (LLMs) behave when their initial answers are challenged by follow-up prompts such as: “Are you sure?” “You are wrong!” confidence elicitation prompts Rather… See the full description on the dataset page: https://huggingface.co/datasets/Reza-Telus/certainty-robustness-llm-evaluation.documentn<1K0 likes52 downloads6mo agoHugging Face28wayne-redemption /Sensor_Driven_Environmental_Monitoring_LLM_Evaluation_Dataset 📌 Dataset Contents Each sample includes: category: The evaluation domain prompt: The question given to the LLM temperature: Environmental temperature input humidity: Environmental humidity input context: A scenario label (e.g., cool_humid, hot_dry, average_day) reference: Expert-crafted expected output All data is provided in a single JSON file. 🧪 Intended Use This dataset supports research on: LLM evaluation methods (semantic similarity, contextual… See the full description on the dataset page: https://huggingface.co/datasets/wayne-redemption/Sensor_Driven_Environmental_Monitoring_LLM_Evaluation_Dataset.tabulartext-classificationn<1K0 likes51 downloads10mo agoHugging Face29AnjanSB /NQ-RAG-DPO-Evaluation Dataset Card Dataset Summary This repository contains five structured data files forming a complete evaluation and training workflow for a multi-perspective alignment system built on Retrieval-Augmented Generation (RAG) and Direct Preference Optimization (DPO). The system is organized into three interconnected pipelines: 1️. RAG Pipeline The RAG pipeline uses the Natural Questions (validation split) as the knowledge source and evaluation benchmark. For each… See the full description on the dataset page: https://huggingface.co/datasets/AnjanSB/NQ-RAG-DPO-Evaluation.texttext-generation1K<n<10K1 likes49 downloads7mo agoHugging Face30Rajan2026 /soas-english-uzbek-rag-evaluation SOAS English-Uzbek Retrieval Pilot Dataset Summary This folder documents a bilingual English-Uzbek retrieval evaluation benchmark for culturally grounded RAG systems. The 400-row public pilot release is retrieval-only: it contains questions and source-document targets, but it intentionally excludes answer, context, excerpt, and source-text fields. This is a pilot benchmark with documented quality flags, template-generated examples, and domain mismatches. The rows… See the full description on the dataset page: https://huggingface.co/datasets/Rajan2026/soas-english-uzbek-rag-evaluation.texttext-retrievaln<1K0 likes49 downloads18d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.