CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CTPLab-DBE-UniBas /staining-robustness-evaluation A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models This repository provides the stain references, pretrained models, and experimental results required to: Define custom staining references using our PLISM reference library Reproduce our published controlled staining robustness experiments 👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main 👉 Associated publication: Paper Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.tabular100K<n<1M1 likes2.7k downloads4mo agoHugging Face02allganize /RAG-Evaluation-Dataset-KO Allganize RAG Leaderboard Allganize RAG 리더보드는 5개 도메인(금융, 공공, 의료, 법률, 커머스)에 대해서 한국어 RAG의 성능을 평가합니다.일반적인 RAG는 간단한 질문에 대해서는 답변을 잘 하지만, 문서의 테이블과 이미지에 대한 질문은 답변을 잘 못합니다. RAG 도입을 원하는 수많은 기업들은 자사에 맞는 도메인, 문서 타입, 질문 형태를 반영한 한국어 RAG 성능표를 원하고 있습니다.평가를 위해서는 공개된 문서와 질문, 답변 같은 데이터 셋이 필요하지만, 자체 구축은 시간과 비용이 많이 드는 일입니다.이제 올거나이즈는 RAG 평가 데이터를 모두 공개합니다. RAG는 Parser, Retrieval, Generation 크게 3가지 파트로 구성되어 있습니다.현재, 공개되어 있는 RAG 리더보드 중, 3가지 파트를 전체적으로 평가하는 한국어로 구성된 리더보드는 없습니다. Allganize RAG 리더보드에서는 문서를… See the full description on the dataset page: https://huggingface.co/datasets/allganize/RAG-Evaluation-Dataset-KO.textn<1K115 likes356 downloads2y agoHugging Face03egolimblevskaia /circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations CircuitLens & WeightLens: Transcoder Descriptions and Evaluations This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods. Methods CircuitLens: https://github.com/egolimblevskaia/CircuitLens WeightLens: https://github.com/egolimblevskaia/WeightLens Dataset Structure The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.tabulartext-classification10K<n<100K0 likes295 downloads7mo agoHugging Face04felixleungsc /paperswithcode-data-evaluation-tables Process data from paperswithcode See https://huggingface.co/datasets/pwc-archive/files/tree/main. Download and unzip evaluation tables: curl -L -O "https://huggingface.co/datasets/pwc-archive/files/resolve/main/jul-28-evaluation-tables.json.gz" gunzip jul-28-evaluation-tables.json.gz Install jq. See https://jqlang.org/. If on Debian/Ubuntu, install with sudo apt-get install jq. Example jq to extract: jq -r ' def process(parent): .task as $current_task | (if parent then… See the full description on the dataset page: https://huggingface.co/datasets/felixleungsc/paperswithcode-data-evaluation-tables.text100K<n<1M1 likes265 downloads11mo agoHugging Face05allganize /RAG-Evaluation-Dataset-JA Allganize RAG Leaderboard とは Allganize RAG Leaderboard は、5つの業種ドメイン(金融、情報通信、製造、公共、流通・小売)において、日本語のRAGの性能評価を実施したものです。一般的なRAGは簡単な質問に対する回答は可能ですが、図表の中に記載されている情報などに対して回答できないケースが多く存在します。RAGの導入を希望する多くの企業は、自社と同じ業種ドメイン、文書タイプ、質問形態を反映した日本語のRAGの性能評価を求めています。RAGの性能評価には、検証ドキュメントや質問と回答といったデータセット、検証環境の構築が必要となりますが、AllganizeではRAGの導入検討の参考にしていただきたく、日本語のRAG性能評価に必要なデータを公開いたしました。RAGソリューションは、Parser、Retrieval、Generation の3つのパートで構成されています。現在、この3つのパートを総合的に評価した日本語のRAG Leaderboardは存在していません。(公開時点)Allganize RAG… See the full description on the dataset page: https://huggingface.co/datasets/allganize/RAG-Evaluation-Dataset-JA.textn<1K34 likes255 downloads2y agoHugging Face06chillies /IELTS-writing-task-2-evaluationtext10K<n<100K39 likes233 downloads3y agoHugging Face07plnguyen2908 /AudioVisual-Benchmark-Evaluation AudioVisual Benchmark Evaluation — evaluation subsets Item-id lists for the audio-visual benchmark subsets used in our reported evaluation tables. Layout <benchmark>/eval_subset.csv item ids evaluated in the paper <benchmark>/media_index.csv id -> media filename(s) <benchmark>/media/ the media files those ids refer to eval_subset.csv holds a single id column keyed to the source benchmark (question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.audiomultiple-choice10K<n<100K0 likes208 downloads22d agoHugging Face08RicardoRei /wmt-da-human-evaluation Dataset Summary This dataset contains all DA human annotations from previous WMT News Translation shared tasks. The data is organised into 8 columns: lp: language pair src: input text mt: translation ref: reference translation score: z score raw: direct assessment annotators: number of annotators domain: domain of the input text (e.g. news) year: collection year You can also find the original data for each year in the results section https://www.statmt.org/wmt{YEAR}/results.html… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-da-human-evaluation.tabular1M<n<10M10 likes205 downloads4y agoHugging Face09RicardoRei /wmt-mqm-human-evaluation Dataset Summary This dataset contains all MQM human annotations from previous WMT Metrics shared tasks and the MQM annotations from Experts, Errors, and Context. The data is organised into 8 columns: lp: language pair src: input text mt: translation ref: reference translation score: MQM score system: MT Engine that produced the translation annotators: number of annotators domain: domain of the input text (e.g. news) year: collection year You can also find the original data here.… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-mqm-human-evaluation.tabular100K<n<1M1 likes185 downloads4y agoHugging Face10TechWolf /JobBERT-evaluation-dataset JobBERT evaluation dataset 💾 This is the official repository containing the evaluation data that was used for the JobBERT paper. This dataset is a list of vacancy titles, each tagged with an ESCO (v1.0.5) occupation. The full dataset is split into two files in a stratified way by class distribution. This data was automatically collected from a large governmental job board. Access the JobBERT paper here: https://arxiv.org/abs/2109.09605 BibTeX Citation If you use this… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/JobBERT-evaluation-dataset.texttext-classification10K<n<100K0 likes146 downloads11mo agoHugging Face11rasinmuhammed /ecommerce-analytics-sql-evaluation Ecommerce Analytics SQL Evaluation (declared GMV, verified answer key) An evalpack: an evaluation database generated from the answer key, not annotated after the fact. A VLDB 2026 audit found 52.8% of BIRD Mini-Dev answer keys wrong because benchmarks annotate answers onto existing databases; this dataset inverts the order. The declared properties (curves, shares, identities) are the specification, the database is generated to satisfy them exactly, and every shipped question was… See the full description on the dataset page: https://huggingface.co/datasets/rasinmuhammed/ecommerce-analytics-sql-evaluation.tabulartable-question-answering10K<n<100K0 likes113 downloads1mo agoHugging Face12rasinmuhammed /saas-finance-sql-evaluation SaaS Finance SQL Evaluation (MRR waterfalls that reconcile exactly) An evalpack: an evaluation database generated from the answer key, not annotated after the fact. A VLDB 2026 audit found 52.8% of BIRD Mini-Dev answer keys wrong because benchmarks annotate answers onto existing databases; this dataset inverts the order. The declared properties (curves, shares, identities) are the specification, the database is generated to satisfy them exactly, and every shipped question was… See the full description on the dataset page: https://huggingface.co/datasets/rasinmuhammed/saas-finance-sql-evaluation.tabulartable-question-answering1K<n<10K0 likes103 downloads1mo agoHugging Face13RicardoRei /wmt-sqm-human-evaluation Dataset Summary In 2022, several changes were made to the annotation procedure used in the WMT Translation task. In contrast to the standard DA (sliding scale from 0-100) used in previous years, in 2022 annotators performed DA+SQM (Direct Assessment + Scalar Quality Metric). In DA+SQM, the annotators still provide a raw score between 0 and 100, but also are presented with seven labeled tick marks. DA+SQM helps to stabilize scores across annotators (as compared to DA). The data is… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-sqm-human-evaluation.tabular100K<n<1M1 likes90 downloads4y agoHugging Face14nwhite-systems /responsible-agent-workflow-evaluation Responsible Agent Workflow Evaluation Version 1.0.0 contains 130 wholly synthetic scenarios for evaluating whether an AI agent respects safety, permission and accountability boundaries in operational settings. Thirteen categories contain ten scenarios each. Every record includes an intentionally unsafe request, contextual facts, expected safe behaviour, explicitly prohibited behaviour, severity, evaluation criteria and reviewer guidance. This is a red-team and… See the full description on the dataset page: https://huggingface.co/datasets/nwhite-systems/responsible-agent-workflow-evaluation.textn<1K1 likes85 downloads2mo agoHugging Face15Cross-Mergeability /extrinsic-evaluations Extrinsic evaluations — the union view One tidy long-format table of every extrinsic (downstream, task-level) evaluation produced across the 2026-08-26 mergeability workstreams, so that a single file answers "how did model X score on benchmark Y" regardless of which experiment produced it. The per-experiment datasets remain the authoritative record of their own methods, figures and caveats. This is the union view, not a replacement, and it deliberately carries no analysis of its… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/extrinsic-evaluations.tabular1K<n<10K0 likes85 downloads27d agoHugging Face16rasinmuhammed /edtech-sql-evaluation EdTech SQL Evaluation (declared learning curve, verified answer key) An evalpack: an evaluation database generated from the answer key, not annotated after the fact. A VLDB 2026 audit found 52.8% of BIRD Mini-Dev answer keys wrong because benchmarks annotate answers onto existing databases; this dataset inverts the order. The declared properties (curves, shares, identities) are the specification, the database is generated to satisfy them exactly, and every shipped question was… See the full description on the dataset page: https://huggingface.co/datasets/rasinmuhammed/edtech-sql-evaluation.tabulartable-question-answering10K<n<100K0 likes74 downloads1mo agoHugging Face17recruit-jp /japanese-image-classification-evaluation-dataset recruit-jp/japanese-image-classification-evaluation-dataset Overview Developed by: Recruit Co., Ltd. Dataset type: Image Classification Language(s): Japanese LICENSE: CC-BY-4.0 More details are described in our tech blog post. 日本語CLIP学習済みモデルとその評価用データセットの公開 Dataset Details This dataset is comprised of four image classification tasks related to concepts and things unique to Japan. Specifically, is consists of the following tasks. jafood101: Image… See the full description on the dataset page: https://huggingface.co/datasets/recruit-jp/japanese-image-classification-evaluation-dataset.imageimage-classification1K<n<10K8 likes67 downloads3y agoHugging Face18WueNLP /mHallucination_Evaluation Multilingual Hallucination Evaluation in the wild The dataset was as part of the paper: How Much Do LLMs Hallucinate across Languages? On Multilingual Estimation of LLM Hallucination in the Wild Below is the figure summarizing the multilingual hallucination evaluation dataset creation (and multilingual hallucination detection dataset): Dataset Details The dataset is a high quality synthetic query/prompt and wikipedia reference pair for estimating hallucinations in the… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/mHallucination_Evaluation.texttext-generation10K<n<100K0 likes62 downloads2y agoHugging Face19nbvbharath-1729 /llm-math-evaluation-dataset LLM Math Response Evaluation Dataset Dataset Summary A human-annotated dataset of 150 AI-generated math responses evaluated across GPT-4o, Claude, and Gemini. Each response is scored on Correctness, Reasoning, and Clarity using a structured rubric, with written justification for every score. Supported Tasks LLM evaluation and benchmarking Math reasoning quality assessment Error type classification in AI responses Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nbvbharath-1729/llm-math-evaluation-dataset.tabularn<1K0 likes62 downloads3d agoHugging Face20patchmedia-org /remote-ai-evaluation-training-market-snapshot Dataset Description This is an aggregate August 22, 2026 research snapshot from Specialist AI Work, an independent PatchMedia tracker of reviewed remote AI evaluation, AI training, data annotation-adjacent, and expert-review opportunities. The live Specialist AI Work inventory has advanced since this snapshot. The counts in this repository describe the immutable August 22 research object; they are not a claim about today's inventory. Reporting date: 2026-08-22 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/patchmedia-org/remote-ai-evaluation-training-market-snapshot.tabularn<1K0 likes60 downloads1mo agoHugging Face21aiagentkarl /agent-evaluation-benchmark Agent Evaluation Benchmark A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories. Overview This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more. Categories Category Test Cases Description Weather 5 Forecasts, UV index, climate history Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.texttext-generationn<1K0 likes57 downloads6mo agoHugging Face22flax-sentence-embeddings /Gender_Bias_Evaluation_SetThis dataset has been created as part of the Flax/JAX community week for testing the flax-sentence-embeddings Sentence Similarity models for Gender Bias but can be used for other use-cases as well related to evaluating Gender Bias. The Following Dataset has been created for Evaluating Gender Bias for different models, based on various stereotypical occupations. The Structure of the dataset is of the following type: Base Sentence Occupation Steretypical_Gender Male Sentence Female… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/Gender_Bias_Evaluation_Set.text1K<n<10K4 likes56 downloads2mo agoHugging Face23OdiaGenAI /RAG_Evaluation_Datasettabular1K<n<10K0 likes51 downloads3y agoHugging Face24g-for-gour /llm-commit-message-evaluation Dataset Card for LLM Commit Message Evaluation The LLM Commit Message Evaluation dataset is designed to evaluate and compare the performance of Large Language Models (LLMs) in generating high-quality git commit messages. It contains real-world code diffs, issue descriptions, and issue titles extracted from open-source repositories (such as OWASP/Nest). For each code change, the dataset provides the original human-written commit message alongside commit messages generated by… See the full description on the dataset page: https://huggingface.co/datasets/g-for-gour/llm-commit-message-evaluation.tabularn<1K0 likes51 downloads2mo agoHugging Face25deutsche-telekom /NLU-Evaluation-Data-en-de NLU Evaluation Data - English and German A labeled English and German language multi-domain dataset (21 domains) with 25K user utterances for human-robot interaction. This dataset is collected and annotated for evaluating NLU services and platforms. The detailed paper on this dataset can be found at arXiv.org: Benchmarking Natural Language Understanding Services for building Conversational Agents The dataset builds on the annotated data of the xliuhw/NLU-Evaluation-Data repository.… See the full description on the dataset page: https://huggingface.co/datasets/deutsche-telekom/NLU-Evaluation-Data-en-de.tabulartext-classification10K<n<100K2 likes47 downloads3y agoHugging Face26awaaz-se-alfaaz /YouTube-Evaluation-Set Awaaz se Alfaaz — YouTube Evaluation Set This dataset is the realistic multi-speaker evaluation set used in Awaaz se Alfaaz, accepted at LaTeLL 2026 — "Enhancing Urdu ASR with Whisper v3: Fine-Tuning on Latest Datasets and Realistic Multi-Speaker Evaluation with SLM Post-Processing." It contains 30 short-form Urdu YouTube videos (YouTube Shorts) covering a mix of news, sports, and current affairs content, along with human annotated gold transcripts and transcripts produced by… See the full description on the dataset page: https://huggingface.co/datasets/awaaz-se-alfaaz/YouTube-Evaluation-Set.textautomatic-speech-recognitionn<1K0 likes42 downloads1mo agoHugging Face27pbevan11 /image_gen_ocr_evaluation_data image_gen_ocr_eval Author: Peter J. Bevan Date: 15/12/23 github: https://github.com/pbevan1/image-gen-spelling-eval Table 1: Normalised Levenshtein similarity scores between instructed text and text present in image (as identified by OCR) Model object signage natural long Overall DALLE3 0.62 0.62 0.62 0.58 0.61 DeepFloydIF 0.57 0.56 0.66 0.39 0.54 DALLE2 0.44 0.35 0.42 0.22 0.36 SDXL 0.3 0.33 0.4 0.21 0.31 SD 0.28 0.26 0.32 0.22 0.27 PlayGroundV2 0.19 0.23 0.17… See the full description on the dataset page: https://huggingface.co/datasets/pbevan11/image_gen_ocr_evaluation_data.textn<1K0 likes36 downloads2y agoHugging Face28ucberkeley-dlab /normative_evaluation_llms_everyday_dilemmastabular10K<n<100K2 likes35 downloads1y agoHugging Face29tuandarcy /IELTS-writing-task-2-evaluationtext10K<n<100K0 likes35 downloads15d agoHugging Face30Equall /perplexity_evaluation SaulLM-7B: Pioneering the first Legal Large Language Model Perplexity Analysis This dataset presents the data used in the paper "SaulLM-7B: Pioneering the first Legal Large Language Model" in "6.3 Perplexity Analysis" section. The dataset contains the perplexity scores of SaulLM-7B, Llama2-7B and Mistral-7B across a corpora of recent text. Cleaning We proceeded to standardize the data by removing any special characters using unicodedata normalization. We also… See the full description on the dataset page: https://huggingface.co/datasets/Equall/perplexity_evaluation.tabular1K<n<10K3 likes31 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.