CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01anon-cmevs-2026 /cmevs-erp-eval CM-EVS: A Coverage-Curated Panoramic RGB-D Dataset for Indoor Scene Understanding CM-EVS is a curated panoramic RGB-D dataset built under a single principle: maximize the geometric coverage of a 3D scene with the fewest equirectangular (ERP) frames possible. The release is structured as one redistributable Blender indoor data archive plus four license-aware adapter packages that regenerate matched frames locally from upstream sources whose terms forbid redistribution. v1.0… See the full description on the dataset page: https://huggingface.co/datasets/anon-cmevs-2026/cmevs-erp-eval.imagedepth-estimationn<1K9 likes21k downloads4mo agoHugging Face02CTPLab-DBE-UniBas /staining-robustness-evaluation A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models This repository provides the stain references, pretrained models, and experimental results required to: Define custom staining references using our PLISM reference library Reproduce our published controlled staining robustness experiments 👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main 👉 Associated publication: Paper Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.tabular100K<n<1M1 likes2.7k downloads4mo agoHugging Face03KRAFTON /Raon-OpenTTS-Eval Raon-OpenTTS-Eval Technical Report A robustness-oriented evaluation benchmark for zero-shot text-to-speech, covering 4 acoustic regimes (Clean, Noisy, Wild, Expressive) across 12 datasets with 6,000 prompt–text pairs. Existing zero-shot TTS benchmarks typically evaluate models using prompts drawn from a single read-speech dataset, providing an incomplete view of robustness under realistic and challenging recording scenarios. Raon-OpenTTS-Eval… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/Raon-OpenTTS-Eval.audiotext-to-speech1K<n<10K9 likes1.9k downloads4mo agoHugging Face04junhee1998 /gr00t-n15-robocasa-gr1-eval GR00T N1.5 on RoboCasa GR-1 Tabletop — Evaluation Trajectories Per-simulator-step recordings of 1,200 evaluation episodes (24 tasks × 50 episodes) of NVIDIA's GR00T N1.5 vision-language-action model on the RoboCasa GR-1 Tabletop Tasks benchmark. Each episode stores every low-level transition — ego-view frames, robot state, executed actions, rewards, full MuJoCo state, and 3D poses of every object and scene body — so that rollouts can be re-analyzed or re-rendered without… See the full description on the dataset page: https://huggingface.co/datasets/junhee1998/gr00t-n15-robocasa-gr1-eval.textrobotics1K<n<10K0 likes1.2k downloads22d agoHugging Face05MVU-Eval-Team /MVU-Eval-Data MVU-Eval Dataset Paper | Code | Project Page Dataset Description The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive benchmark… See the full description on the dataset page: https://huggingface.co/datasets/MVU-Eval-Team/MVU-Eval-Data.tabularvideo-text-to-text1K<n<10K2 likes1.1k downloads11mo agoHugging Face06evaluate /imdb-citextn<1K0 likes969 downloads4y agoHugging Face07ckadirt /mev1_evalstextn<1K0 likes918 downloads3y agoHugging Face08FerrariKazu /rhan-eval-sweeptabularn<1K0 likes825 downloads13m agoHugging Face09nvidia /PhysicalAI-Robotics-PhysicalAssets-VoMP-Eval VoMP: Predicting Volumetric Mechanical Properties Dataset Description: The Pre-Processed 3D Dataset is a dataset that is composed of 4 individual 3D asset datasets which are processed to render them from multiple views, voxelize the assets, and propagate VLM annotations for material properties. We release pre-processed data derived from the 3D assets, specifically: voxels, rendered images, and LLM-annotated material descriptions. This dataset is for research and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-PhysicalAssets-VoMP-Eval.text1K<n<10K3 likes678 downloads8mo agoHugging Face10AvoCahDoe /llava-15-rlmpq-vlm-eval-results RL-MPQ VLM Evaluation Artifacts Complete figures, tables, galleries, and raw benchmark CSVs for the extended VLM evaluation. Dataset: AvoCahDoe/llava-15-rlmpq-vlm-eval-results Collections (by base VLM) RL-MPQ VLM — LLaVA-1.5-13B — HF collection RL-MPQ VLM — LLaVA-1.5-7B — HF collection RL-MPQ VLM — LLaVA-Next Mistral-7B — HF collection RL-MPQ VLM — Qwen2-VL-7B — HF collection Model repos RL-MPQ High Fidelity →… See the full description on the dataset page: https://huggingface.co/datasets/AvoCahDoe/llava-15-rlmpq-vlm-eval-results.imagevisual-question-answeringn<1K0 likes521 downloads3mo agoHugging Face11jang1563 /narrow-model-safety-eval Narrow Model Safety Evaluation — Protein Dual-Use Risk Dataset Summary: Annotations, results, and evaluation data for a proof-of-concept framework assessing dual-use risk in narrow scientific AI models. Two lines of work: (1) structure-level metrics — FSPE, FSI, and Physical Realizability Tier — on eight published protein toxins and mechanism-matched benign controls (ESM-2, ProteinMPNN); (2) mechanism generalization — a leave-one-mechanism-out panel measuring what an… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/narrow-model-safety-eval.tabularothern<1K0 likes491 downloads3d agoHugging Face12allenai /tulu-3-harmbench-evalThis data comes from the HarmBench benchmark. This is one of the datasets included in the Ai2 Safety Evaluation Suite, and the Tülu 3 evaluation suite. The repo for Ai2's safety suite includes instructions on how to evaluate models on various safety-related evaluation including this one. textn<1K3 likes455 downloads1y agoHugging Face13stephenbasd /MVU-Eval-Data MVU-Eval Dataset Paper | Code | Project Page Dataset Description The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/stephenbasd/MVU-Eval-Data.tabularvideo-text-to-text1K<n<10K0 likes446 downloads3mo agoHugging Face14allganize /RAG-Evaluation-Dataset-KO Allganize RAG Leaderboard Allganize RAG 리더보드는 5개 도메인(금융, 공공, 의료, 법률, 커머스)에 대해서 한국어 RAG의 성능을 평가합니다.일반적인 RAG는 간단한 질문에 대해서는 답변을 잘 하지만, 문서의 테이블과 이미지에 대한 질문은 답변을 잘 못합니다. RAG 도입을 원하는 수많은 기업들은 자사에 맞는 도메인, 문서 타입, 질문 형태를 반영한 한국어 RAG 성능표를 원하고 있습니다.평가를 위해서는 공개된 문서와 질문, 답변 같은 데이터 셋이 필요하지만, 자체 구축은 시간과 비용이 많이 드는 일입니다.이제 올거나이즈는 RAG 평가 데이터를 모두 공개합니다. RAG는 Parser, Retrieval, Generation 크게 3가지 파트로 구성되어 있습니다.현재, 공개되어 있는 RAG 리더보드 중, 3가지 파트를 전체적으로 평가하는 한국어로 구성된 리더보드는 없습니다. Allganize RAG 리더보드에서는 문서를… See the full description on the dataset page: https://huggingface.co/datasets/allganize/RAG-Evaluation-Dataset-KO.textn<1K115 likes356 downloads2y agoHugging Face15autoiac-project /iac-eval IaC-Eval dataset (v1.1) IaC-Eval dataset is the first human-curated and challenging Cloud Infrastructure-as-Code (IaC) dataset tailored to more rigorously benchmark large language models' IaC code generation capabilities. This dataset contains 458 questions ranging from simple to difficult across various cloud services (targeting AWS for now). | Github | 🏆 Leaderboard TBD | 📖 NeurIPS 2024 Paper | 2. Usage instructions Option 1: Running the evaluation… See the full description on the dataset page: https://huggingface.co/datasets/autoiac-project/iac-eval.texttext-generationn<1K7 likes323 downloads2y agoHugging Face16UW-Madison-Lee-Lab /MMLU-Pro-CoT-Eval Dataset Details Modality: Text Format: CSV Size: 100K - 1M rows Total Rows: 248,836 License: MIT Libraries Supported: datasets, pandas, croissant Structure Each row in the dataset includes: question: The query posed in the dataset. answer: The correct response. category: The domain of the question (e.g., math, science). src: The source of the question. id: A unique identifier for each entry. chain_of_thoughts: Step-by-step reasoning steps leading to the answer.… See the full description on the dataset page: https://huggingface.co/datasets/UW-Madison-Lee-Lab/MMLU-Pro-CoT-Eval.text100K<n<1M0 likes308 downloads2y agoHugging Face17Cross-Mergeability /beetle-merge-eval Beetle merged models — benchmark evaluation against their parents Minimal-pair benchmark accuracy for the Beetle merged models published in the Mergeability org, scored against their own parent models and, where one exists, the jointly-trained ceiling on the same harness. The existing sweep datasets (Mergeability/merge-sweep-results, Mergeability-2/mergeability-results) record merge quality in nats (NLL, delta_floor, rel_damage, barrier, geometry). They contain no downstream… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/beetle-merge-eval.tabular1K<n<10K0 likes305 downloads27d agoHugging Face18egolimblevskaia /circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations CircuitLens & WeightLens: Transcoder Descriptions and Evaluations This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods. Methods CircuitLens: https://github.com/egolimblevskaia/CircuitLens WeightLens: https://github.com/egolimblevskaia/WeightLens Dataset Structure The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.tabulartext-classification10K<n<100K0 likes295 downloads7mo agoHugging Face19felixleungsc /paperswithcode-data-evaluation-tables Process data from paperswithcode See https://huggingface.co/datasets/pwc-archive/files/tree/main. Download and unzip evaluation tables: curl -L -O "https://huggingface.co/datasets/pwc-archive/files/resolve/main/jul-28-evaluation-tables.json.gz" gunzip jul-28-evaluation-tables.json.gz Install jq. See https://jqlang.org/. If on Debian/Ubuntu, install with sudo apt-get install jq. Example jq to extract: jq -r ' def process(parent): .task as $current_task | (if parent then… See the full description on the dataset page: https://huggingface.co/datasets/felixleungsc/paperswithcode-data-evaluation-tables.text100K<n<1M1 likes265 downloads11mo agoHugging Face20allganize /RAG-Evaluation-Dataset-JA Allganize RAG Leaderboard とは Allganize RAG Leaderboard は、5つの業種ドメイン(金融、情報通信、製造、公共、流通・小売)において、日本語のRAGの性能評価を実施したものです。一般的なRAGは簡単な質問に対する回答は可能ですが、図表の中に記載されている情報などに対して回答できないケースが多く存在します。RAGの導入を希望する多くの企業は、自社と同じ業種ドメイン、文書タイプ、質問形態を反映した日本語のRAGの性能評価を求めています。RAGの性能評価には、検証ドキュメントや質問と回答といったデータセット、検証環境の構築が必要となりますが、AllganizeではRAGの導入検討の参考にしていただきたく、日本語のRAG性能評価に必要なデータを公開いたしました。RAGソリューションは、Parser、Retrieval、Generation の3つのパートで構成されています。現在、この3つのパートを総合的に評価した日本語のRAG Leaderboardは存在していません。(公開時点)Allganize RAG… See the full description on the dataset page: https://huggingface.co/datasets/allganize/RAG-Evaluation-Dataset-JA.textn<1K34 likes255 downloads2y agoHugging Face21chillies /IELTS-writing-task-2-evaluationtext10K<n<100K39 likes233 downloads3y agoHugging Face22iarfmoose /qa_evaluatorThis is the same dataset as the question_generator dataset but with the context removed and the question and answer in separate fields. This is intended to be used with the question_generator repo to train the qa_evaluator model which predicts whether a question and answer pair makes sense. text100K<n<1M4 likes219 downloads5y agoHugging Face23HAERAE-HUB /K2-EvalResearch Paper coming soon! K2EvalK^{2} EvalK2Eval K2EvalK^{2} EvalK2Eval is a novel benchmark featuring 90 handwritten instructions that require in-depth knowledge of Korean language and culture for accurate completion. Benchmark Overview The design principle behind K2EvalK^{2} EvalK2Eval centers on collecting instructions that necessitate knowledge specific to Korean culture and context in order to solve. This approach distinguishes our work from simply translating… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/K2-Eval.textn<1K8 likes212 downloads2y agoHugging Face24allenai /olmo-eval-strongrejectThis data comes from the StrongREJECT benchmark. This is one of the datasets included in the Ai2 Safety Evaluation Suite, and the Olmo evaluation suite. The repo for Ai2's safety suite includes instructions on how to evaluate models on various safety-related evaluations, including this one. Permitted Use The data is provided for benchmarking and evaluation purposes only. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. Disclaimer This… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmo-eval-strongreject.text10K<n<100K1 likes211 downloads2mo agoHugging Face25plnguyen2908 /AudioVisual-Benchmark-Evaluation AudioVisual Benchmark Evaluation — evaluation subsets Item-id lists for the audio-visual benchmark subsets used in our reported evaluation tables. Layout <benchmark>/eval_subset.csv item ids evaluated in the paper <benchmark>/media_index.csv id -> media filename(s) <benchmark>/media/ the media files those ids refer to eval_subset.csv holds a single id column keyed to the source benchmark (question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.audiomultiple-choice10K<n<100K0 likes208 downloads22d agoHugging Face26reilleo /new_justify_eval_actstextn<1K0 likes206 downloads11mo agoHugging Face27RicardoRei /wmt-da-human-evaluation Dataset Summary This dataset contains all DA human annotations from previous WMT News Translation shared tasks. The data is organised into 8 columns: lp: language pair src: input text mt: translation ref: reference translation score: z score raw: direct assessment annotators: number of annotators domain: domain of the input text (e.g. news) year: collection year You can also find the original data for each year in the results section https://www.statmt.org/wmt{YEAR}/results.html… See the full description on the dataset page: https://huggingface.co/datasets/RicardoRei/wmt-da-human-evaluation.tabular1M<n<10M10 likes205 downloads4y agoHugging Face28previtus /STARCOP_allbands_Evalgated STARCOP dataset STARCOP dataset: Semantic Segmentation of Methane Plumes with Hyperspectral Machine Learning Models 🌈🛰️Authors: Vít Růžička, Gonzalo Mateo-Garcia, Luis Gómez-Chova, Anna Vaughan, Luis Guanter and Andrew Markham Fast data preview in: dataset_exploration.ipynb Main repository: github/spaceml-org/STARCOP Please refer to the main dataset readme file on: https://huggingface.co/datasets/previtus/STARCOP_allbands_Train1 imageimage-segmentation100K<n<1M1 likes198 downloads2y agoHugging Face29baryonlabs /open-ko-s2s-eval-artifacts Open Ko-S2S 평가 산출물 (감사용) ⚠️ KsponSpeech 참조 전사는 해시로 대체돼 있습니다 KsponSpeech 는 AI Hub 배포 데이터로 재배포 제한이 있을 수 있어, kspon 런의 ref 컬럼을 ref_sha256 으로 대체했습니다(전사 원문 미포함). 모델 출력(hyp)과 채점 결과(cer_err/cer_len/cer)는 우리 산출물이라 그대로 공개합니다. Zeroth 런은 원본이 CC BY 4.0(OpenSLR #40)이라 ref 원문을 그대로 담고 있습니다. 라이선스 보유자의 검증 절차 AI Hub 에서 KsponSpeech 를 정당하게 받은 분은 다음으로 우리 수치를 검증할 수 있습니다. 리더보드 저장소의 eval/datasets_ko.py 에서 clean_kspon() 을 가져옵니다. 자기 사본의 원 전사에 clean_kspon() 을 적용합니다. 결과가 목록이면… See the full description on the dataset page: https://huggingface.co/datasets/baryonlabs/open-ko-s2s-eval-artifacts.tabularautomatic-speech-recognitionn<1K0 likes190 downloads1mo agoHugging Face30cruciverb-it /evalita2026 This repository contains the data release for the Cruciverb-IT shared task on automatic crossword solving in Italian, as part of the 2026 EVALITA campaign. Refer to the task website for more details. The data from both tasks can be downloaded from the 'Files and versions' tab. Updates: Minor update to both task_*_scorer.py in order to convert accented letters to their non-accented counterpart during evaluation Test data is out!! The test data of both… See the full description on the dataset page: https://huggingface.co/datasets/cruciverb-it/evalita2026.texttext-generation100K<n<1M4 likes188 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.