CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Rapidata /svg-benchmark Rapidata Static SVG Generation Benchmark Built by Rapidata. This dataset contains 1,918,367 human responses, collected with the Rapidata Python SDK, comparing how well 42 frontier LLMs generate static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment). The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/svg-benchmark.imagetext-to-image100K<n<1M35 likes9.9k downloads1mo agoHugging Face02TheKernel01 /AIGC-Detection-Benchmark AIGC Detection Benchmark Dataset 📝 Dataset Description Dataset Summary The AIGC Detection Benchmark Dataset is a high-quality collection of images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. The dataset contains a mix of real-world images and images generated by a wide array of prominent AI models, including diffusion models (like Stable Diffusion, DALL-E 2, Midjourney, ADM) and GANs… See the full description on the dataset page: https://huggingface.co/datasets/TheKernel01/AIGC-Detection-Benchmark.imageimage-classification100K<n<1M0 likes1.6k downloads6mo agoHugging Face03etri-vilab /holisafe-benchgated ⚠️ CONTENT WARNING: This dataset contains potentially harmful and sensitive visual content including violence, hate speech, illegal activities, self-harm, sexual content, and other unsafe materials. Images are intended solely for safety research and evaluation purposes. Viewer discretion is strongly advised. HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model (CVPR'26 Findings) 🌐 Website | 📑 Paper 📋 HoliSafe-Bench Dataset… See the full description on the dataset page: https://huggingface.co/datasets/etri-vilab/holisafe-bench.imagevisual-question-answering1K<n<10K11 likes667 downloads4mo agoHugging Face04PCF-Bench /PCF-Bench PCF-Bench A photonic-crystal-fiber (PCF) inverse-design benchmark for vision-language models. Each sample bundles geometric parameters, simulated mode-field images, and a four-level expert-style annotation suite supporting tasks across geometry perception, physics understanding, multimodal reasoning, inverse design, and code generation. Note (anonymous review). This dataset card omits identifying information for double-blind review. Note (review subset). The full source corpus is… See the full description on the dataset page: https://huggingface.co/datasets/PCF-Bench/PCF-Bench.imagevisual-question-answering10K<n<100K0 likes354 downloads5mo agoHugging Face05NoeFontana /locus-tag-bench Locus-Tag Bench Synthetic benchmark suite for fiducial-tag detection and camera calibration, rendered with render-tag. Each config corresponds to one campaign (board family × resolution, or a lighting/sensor variant). Images are rendered in Blender Cycles with ground-truth geometry recovered directly from the scene — no detector-in-the-loop, no human labels. Configs Config Purpose Images Board/Tag Resolution locus_v1_tag36h11_640x480 Detection (low-res)… See the full description on the dataset page: https://huggingface.co/datasets/NoeFontana/locus-tag-bench.imageobject-detectionn<1K1 likes317 downloads5mo agoHugging Face06MTSAIR /MWS-Antifraud-Bench MWS Antifraud Bench (Validation) Experimental document-authenticity task for general-purpose multimodal language models. This is the public validation part of MWS Vision Bench anti-fraud v0.1. The dataset is released for research and model comparison. It is not a certification tool, a production fraud-detection system, or a universal leaderboard that is expected to be resistant to deliberate optimization. Data The validation split contains 209 items: 44 ai_gen;… See the full description on the dataset page: https://huggingface.co/datasets/MTSAIR/MWS-Antifraud-Bench.imageimage-classificationn<1K0 likes311 downloads2mo agoHugging Face07AI-EcoNet /HUGO-Bench-Paper-Reproducibility HUGO-Bench Paper Reproducibility Supplementary data and reproducibility materials for the paper: Vision Transformers for Zero-Shot Clustering of Animal Images: A Comparative Benchmarking Study - https://arxiv.org/abs/2602.03894 Hugo Markoff, Stefan Hein Bengtson, Michael Ørsted Aalborg University, Denmark Dataset Description This repository contains complete experimental results, pre-computed embeddings, and execution logs from our comprehensive benchmarking study… See the full description on the dataset page: https://huggingface.co/datasets/AI-EcoNet/HUGO-Bench-Paper-Reproducibility.tabularimage-classification100K<n<1M0 likes300 downloads6mo agoHugging Face08FForty7 /svg-benchmark Rapidata Static SVG Generation Benchmark Built by Rapidata. This dataset contains 1,355,161 human responses, collected with the Rapidata Python SDK, comparing how well 30 frontier LLMs generate static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment). The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/FForty7/svg-benchmark.imagetext-to-image100K<n<1M0 likes293 downloads3mo agoHugging Face09MM-Hallu /MFC-Bench MFC-Bench: Multimodal Fact-Checking Benchmark MFC-Bench is a comprehensive Multimodal Fact-Checking testbed designed to evaluate LVLMs in terms of identifying factual inconsistencies and counterfactual scenarios. Dataset Description From the paper: "MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models" MFC-Bench encompasses a wide range of visual and textual queries, organized into three binary classification tasks: 1. Manipulation… See the full description on the dataset page: https://huggingface.co/datasets/MM-Hallu/MFC-Bench.imageimage-classification10K<n<100K1 likes206 downloads5mo agoHugging Face10AI-EcoNet /HUGO-Bench HUGO-Bench Hierarchical Unsupervised Grouping of Organisms Benchmark A comprehensive benchmark dataset for evaluating zero-shot clustering of wildlife camera trap images using Vision Transformer embeddings. Overview HUGO-Bench contains 139,111 expert-validated cropped images of 60 animal species (30 birds, 30 mammals), derived from 23 camera trap projects across LILA BC. The dataset enables benchmarking of Vision Transformer models for unsupervised species-level… See the full description on the dataset page: https://huggingface.co/datasets/AI-EcoNet/HUGO-Bench.imageimage-classification100K<n<1M1 likes186 downloads6mo agoHugging Face11ucsahin /Turkish-VLM-Mix-BenchmarkThis is a Turkish multimodal (image-text-text triplets) dataset consisting of Turkish translated samples from the datasets google/docci, tomg-group-umd/pixelprose, detection-datasets/coco, rafaelpadilla/coco2017, liuhaotian/LLaVA-Instruct-150K, liuhaotian/LLaVA-CC3M-Pretrain-595K, and HuggingFaceM4/FairFace. The labels are in Turkish and the dataset is in an instruction-tuning format with separate columns for prompts and completion labels. The original labels (except… See the full description on the dataset page: https://huggingface.co/datasets/ucsahin/Turkish-VLM-Mix-Benchmark.imageimage-to-text10K<n<100K7 likes143 downloads2y agoHugging Face12nutrientdocs /doc-split-benchmark Doc-Split Benchmark The evaluation slice for page-stream segmentation — the exact set behind the leaderboard and the cloud-VLM comparison. Self-contained (page images embedded), with a reference scorer so results are reproducible. This is the benchmark, not the training corpus (which stays private). 🏆 Leaderboard: doc-split-leaderboard 🎯 Demo: doc-split-demo 🟢 Model: doc-split-mini-e5 (open weights) 🌍 OpenPSS cuts: openpss-mirror (SHORT/LONG, self-contained)… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/doc-split-benchmark.imageimage-classificationn<1K0 likes85 downloads2mo agoHugging Face13act13 /AVA-Bench AVA-Bench Training dataset for the paper AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models (arXiv:2506.09082) accepted in CVPR 2026. AVA-Bench is a diagnostic benchmark for evaluating Vision Foundation Models (VFMs) through Atomic Visual Abilities (AVAs): fundamental perceptual skills such as localization, counting, OCR, spatial understanding, depth estimation, color recognition, texture recognition, and fine-grained recognition. AVA-Bench disentangls visual… See the full description on the dataset page: https://huggingface.co/datasets/act13/AVA-Bench.imagevisual-question-answering100K<n<1M0 likes84 downloads4mo agoHugging Face14thoughtworks /document-processing-benchmark Document Processing Benchmark 8 public document datasets (receipts, invoices, forms, bank statements, multi-page docs, contracts) normalized into one parquet schema. Each row has the document, ground-truth annotations, and per-row token/latency/cost numbers from real API calls to one or more reference models. You can read off a target's cost/latency/quality without re-running it. from datasets import load_dataset ds = load_dataset("thoughtworks/document-processing-benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/document-processing-benchmark.tabularimage-to-text10K<n<100K1 likes73 downloads4mo agoHugging Face15NeeyuHuynh /chest-bench-example ChestBench Example DICOM-VLM Framework Reference Package v0.2.0 ChestBench Example is a four-case, DICOM-native reference package for developing and validating the data architecture of a medical vision-language model (VLM) pipeline. It is intentionally small. Its purpose is to demonstrate how medical imaging data, annotations, text, knowledge, retrieval targets, QA, evidence requirements, perturbations, and audit metadata can be represented without confusing… See the full description on the dataset page: https://huggingface.co/datasets/NeeyuHuynh/chest-bench-example.imageimage-classificationn<1K0 likes72 downloads17d agoHugging Face16thu-coai /Syncred-Bench Syncred-Bench SynCred-Bench is a benchmark designed to evaluate synthetic credibility: AI-generated images that appear trustworthy by imitating authoritative visual forms (e.g., fake notices, credentials, news layouts) and realistic circulation traces. The benchmark contains 600 AI-generated misinformation images across six credible-form categories and seven circulation styles. It also introduces FP450, a real-image negative set for measuring false positives in detection… See the full description on the dataset page: https://huggingface.co/datasets/thu-coai/Syncred-Bench.imageimage-classification1K<n<10K2 likes71 downloads4mo agoHugging Face17nutrientdocs /document-classification-benchmark Document Classification Benchmark (open-vocab, zero-shot) Given a document image and an arbitrary set of text labels, which one is right? A held-out, zero-shot, open-vocabulary evaluation for document-type classification — labels are supplied at inference, not baked into a head. Test split only; not for training. Every image is drawn from a permissively-licensed, redistributable source. Powers the document-classification-leaderboard and evaluates document-classification-v2… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/document-classification-benchmark.imagezero-shot-image-classification1K<n<10K0 likes67 downloads2mo agoHugging Face18dellacorte /PANDA-PLUS-Bench PANDA-PLUS-Bench A benchmark dataset for evaluating WSI-specific feature collapse in pathology foundation models. Dataset Description PANDA-PLUS-Bench contains expert-annotated prostate biopsy patches from 9 whole slide images (9 unique patients) with pixel-level Gleason pattern annotations. Dataset Summary Patches: ~2,770 per augmentation condition Resolution: 224×224 pixels at 20× magnification Classes: Benign (0), GP3 (1), GP4 (2), GP5 (3) Slides: 9 (one… See the full description on the dataset page: https://huggingface.co/datasets/dellacorte/PANDA-PLUS-Bench.imageimage-classification10K<n<100K1 likes62 downloads9mo agoHugging Face19whatfontis /WhatFontIs-Bench WhatFontIs-Bench - A Synthetic Benchmark for Font Family Identification A synthetic test set for font family identification: a single word, set in a known font, printed or painted on real surfaces and in real scenes, with the exact font, the text and the position of every letter recorded for each image. Made by WhatFontIs, the font finder that identifies fonts from images, to measure how well a tool can find the font in a real-looking photo. Official page:… See the full description on the dataset page: https://huggingface.co/datasets/whatfontis/WhatFontIs-Bench.imageimage-classification10K<n<100K0 likes61 downloads5d agoHugging Face20gigzjl /design-fto-bench PatSnap Design FTO Bench A Bench for evaluating design patent Freedom-To-Operate (FTO) retrieval systems on cross-modal image search. Each sample provides a query product image (or design patent figure) plus the ground truth set of target design patents that constitute infringement risk, as confirmed by patent invalidation proceedings. 🐙 GitHub mirror: This dataset is also published as part of the patsnap/patent-bench monorepo, where you can find the reference metric scripts… See the full description on the dataset page: https://huggingface.co/datasets/gigzjl/design-fto-bench.imageimage-to-imagen<1K0 likes55 downloads25d agoHugging Face21Robo531 /ai-detector-benchmark-test-data 🎯 AI Detector Benchmark Test Dataset A comprehensive benchmark dataset for testing AI image detection models. 📊 Dataset Summary Total Images: 700 AI-Generated: 250 images (from 5 different generators) Real Images: 450 images (from 9 diverse datasets) Perfect for: ✅ Testing AI detection models ✅ Creating leaderboards ✅ Comparing model performance ✅ Benchmarking new approaches 🤖 AI Generators Included Generator Images Accuracy Baseline FLUX… See the full description on the dataset page: https://huggingface.co/datasets/Robo531/ai-detector-benchmark-test-data.imageimage-classificationn<1K0 likes50 downloads6mo agoHugging Face22EDAnonSubmission /benchmark EditJudge-Bench EditJudge-Bench is a synthetic benchmark for auditing vision-language models used as automated judges for image-edit verification. Each row contains a source image, an edited image, a factual edit instruction, counterfactual instructions, and ground-truth scene parameters produced by a controlled Blender/Infinigen generation pipeline. This repository is an anonymous review release for a NeurIPS Evaluations and Datasets submission. Dataset Contents 1… See the full description on the dataset page: https://huggingface.co/datasets/EDAnonSubmission/benchmark.imageimage-classification1K<n<10K0 likes48 downloads5mo agoHugging Face23PatSnap /design-fto-bench PatSnap Design FTO Bench A Bench for evaluating design patent Freedom-To-Operate (FTO) retrieval systems on cross-modal image search. Each sample provides a query product image (or design patent figure) plus the ground truth set of target design patents that constitute infringement risk, as confirmed by patent invalidation proceedings. 🐙 GitHub mirror: This dataset is also published as part of the patsnap/patent-bench monorepo, where you can find the reference metric scripts… See the full description on the dataset page: https://huggingface.co/datasets/PatSnap/design-fto-bench.imageimage-to-imagen<1K4 likes48 downloads4mo agoHugging Face24macular /diabetic-retinopathy-screening-benchmark-africa DR-Africa-Benchmark — Screening-Prevalence-Corrected, Fairness-Instrumented DR Evaluation An evaluation benchmark for diabetic-retinopathy grading under African screening conditions. It does not introduce new labels; it introduces evaluation validity — per-record importance weights that reweight a referral-skewed image set to real Sub-Saharan-Africa population prevalence, plus synthetic subgroup metadata for fairness reporting. Version 1.0.0 · core dr_synth 1.0.0 · part of the… See the full description on the dataset page: https://huggingface.co/datasets/macular/diabetic-retinopathy-screening-benchmark-africa.imageimage-classification1K<n<10K0 likes47 downloads4mo agoHugging Face25Rapidata /Face_Generation_Benchmark Rapidata Human Face Generation Alignment This T2I dataset contains over ~22'000 human responses, collected in less than 1h using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation. Evaluating 12 different image generation models on which one can generate faces more accurately. The question that the annotators get asked is: "Which Image follows the description of the human better?" To evaluate your own models and create leaderboard check out our… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Face_Generation_Benchmark.imagetext-to-image1K<n<10K16 likes44 downloads11mo agoHugging Face26nutrientdocs /doc-openvocab-benchmark Open-Vocab Document & Figure Classification Benchmark Given a document or figure image and an arbitrary set of text labels, which one is right? This is a zero-shot, open-vocabulary image-classification benchmark for the document-AI setting: every image is scored against a broad ~48-label candidate vocabulary (document types + figure/zone types), and the task is to pick the correct label. The labels are supplied at inference — which is precisely what a fixed-label supervised… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/doc-openvocab-benchmark.imagezero-shot-image-classification1K<n<10K1 likes41 downloads2mo agoHugging Face27MING-ZCH /TFQ-Bench-Full TFQ-Bench: A Benchmark for Evaluating Image Implication Understanding TFQ-Bench is a rigorous evaluation benchmark designed to assess the capabilities of MLLMs in understanding visual metaphors, sarcasm, and implicit meanings via True-False Questions. It serves as a complement to existing benchmarks like II-Bench (Multiple-Choice Question) and CII-Bench (Open-Style Question), offering a lower-bound difficulty check that tests a model's ability to verify specific propositions about… See the full description on the dataset page: https://huggingface.co/datasets/MING-ZCH/TFQ-Bench-Full.imagevisual-question-answering10K<n<100K1 likes34 downloads8mo agoHugging Face28ash12321 /ai-detector-benchmark-test-data 🎯 AI Detector Benchmark Test Dataset A comprehensive benchmark dataset for testing AI image detection models. 📊 Dataset Summary Total Images: 700 AI-Generated: 250 images (from 5 different generators) Real Images: 450 images (from 9 diverse datasets) Perfect for: ✅ Testing AI detection models ✅ Creating leaderboards ✅ Comparing model performance ✅ Benchmarking new approaches 🤖 AI Generators Included Generator Images Accuracy Baseline FLUX… See the full description on the dataset page: https://huggingface.co/datasets/ash12321/ai-detector-benchmark-test-data.imageimage-classificationn<1K0 likes27 downloads9mo agoHugging Face29BDRC /tibetan-script-classification-benchmark Tibetan Script Classification Benchmark Holdout benchmark for 6-class Tibetan script classification. Test split only — not used during training. All images are BDRC manuscript page scans, balanced by subclass. Class Images Subclasses Danyig 60 DraDring: 25, DraRing: 9, Drathung: 17, Gongshabma: 3, Tsegdrig: 6 Druma 60 Dhumri: 22, DruDring: 20, DruRing: 10, Druchen: 2, Druthung: 6 Gyuyig 60 Khyuyig: 31, Tsumachug: 15, Yigchung: 14 Pedri 60 Peri: 44, Petsuk: 16… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-script-classification-benchmark.imageimage-classificationn<1K0 likes25 downloads3mo agoHugging Face30MING-ZCH /TFQ-Bench-Lite TFQ-Bench: A Benchmark for Evaluating Image Implication Understanding TFQ-Bench is a rigorous evaluation benchmark designed to assess the capabilities of MLLMs in understanding visual metaphors, sarcasm, and implicit meanings via True-False Questions. It serves as a complement to existing benchmarks like II-Bench (Multiple-Choice Question) and CII-Bench (Open-Style Question), offering a lower-bound difficulty check that tests a model's ability to verify specific propositions about… See the full description on the dataset page: https://huggingface.co/datasets/MING-ZCH/TFQ-Bench-Lite.imagevisual-question-answeringn<1K1 likes18 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.