CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Rapidata /svg-benchmark Rapidata Static SVG Generation Benchmark Built by Rapidata. This dataset contains 1,918,367 human responses, collected with the Rapidata Python SDK, comparing how well 42 frontier LLMs generate static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment). The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/svg-benchmark.imagetext-to-image100K<n<1M34 likes10k downloads27d agoHugging Face02gently-project /gently-perception-benchmark Gently Perception Agent Benchmark Light-sheet microscopy volumes of C. elegans embryo development, intended for evaluating vision-based perception agents on embryo stage classification. The dataset has two tiers: Annotated benchmark set (embryo_1–embryo_8) — human ground-truth stage transitions. Use this for evaluation. Unannotated corpus (embryo_9–embryo_105) — 97 additional real embryo timelapses with no human labels, provided for developing and stress- testing perception… See the full description on the dataset page: https://huggingface.co/datasets/gently-project/gently-perception-benchmark.imageimage-classification10K<n<100K1 likes2.1k downloads3mo agoHugging Face03TheKernel01 /AIGC-Detection-Benchmark AIGC Detection Benchmark Dataset 📝 Dataset Description Dataset Summary The AIGC Detection Benchmark Dataset is a high-quality collection of images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. The dataset contains a mix of real-world images and images generated by a wide array of prominent AI models, including diffusion models (like Stable Diffusion, DALL-E 2, Midjourney, ADM) and GANs… See the full description on the dataset page: https://huggingface.co/datasets/TheKernel01/AIGC-Detection-Benchmark.imageimage-classification100K<n<1M0 likes1.1k downloads6mo agoHugging Face04Hammadhaideerr /CTTA-AD-Benchmarks CTTA-AD Benchmarks Dataset collection for CTTA-AD: Continual Test-Time Adaptation for Unified Few-Shot Visual Anomaly Detection (AAAI 2027 submission). Datasets Dataset Domain Categories Train Normal License MVTec-AD Industrial 15 209–391 per category CC BY-NC-SA 4.0 VisA Industrial 12 400–905 per category CC BY-NC-SA 4.0 MVTec-LOCO Logical 5 varies CC BY-NC-SA 4.0 BrainMRI Medical 1 7,500 Research only LiverCT Medical 1 1,542 Research only… See the full description on the dataset page: https://huggingface.co/datasets/Hammadhaideerr/CTTA-AD-Benchmarks.imageimage-classification10K<n<100K0 likes724 downloads4mo agoHugging Face05FForty7 /svg-benchmark Rapidata Static SVG Generation Benchmark Built by Rapidata. This dataset contains 1,355,161 human responses, collected with the Rapidata Python SDK, comparing how well 30 frontier LLMs generate static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment). The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/FForty7/svg-benchmark.imagetext-to-image100K<n<1M0 likes307 downloads3mo agoHugging Face06khadijah00 /ppe-benchmark-eval PPE Benchmark Eval Set (v1) A held-out, human-verified benchmark for evaluating vision-language models on personal protective equipment (PPE) detection — specifically hardhat and safety-vest presence — framed as a VQA-style classification task. What this is 96 images, balanced 24/24/24/24 across the four hardhat × vest combinations (yes/yes, yes/no, no/yes, no/no). Sourced from a forked, filtered subset of the karabuk-university PPE dataset on Roboflow Universe… See the full description on the dataset page: https://huggingface.co/datasets/khadijah00/ppe-benchmark-eval.imagevisual-question-answeringn<1K0 likes167 downloads1mo agoHugging Face07kenobi /GeneLab_BPS_BenchmarkData Dataset Card for Dataset GeneLab_BPS_BenchmarkData Dataset Details This dataset is a version of the Biological and Physical Sciences (BPS) Microscopy Benchmark Training Dataset managed by NASA and hosted on an S3 Bucket here: https://registry.opendata.aws/bps_microscopy/ Fluorescence microscopy images of individual nuclei from mouse fibroblast cells, irradiated with Fe particles or X-rays with fluorescent foci indicating 53BP1 positivity, a marker of DNA damage.… See the full description on the dataset page: https://huggingface.co/datasets/kenobi/GeneLab_BPS_BenchmarkData.imageimage-classification10K<n<100K0 likes129 downloads3y agoHugging Face08ucsahin /Turkish-VLM-Mix-BenchmarkThis is a Turkish multimodal (image-text-text triplets) dataset consisting of Turkish translated samples from the datasets google/docci, tomg-group-umd/pixelprose, detection-datasets/coco, rafaelpadilla/coco2017, liuhaotian/LLaVA-Instruct-150K, liuhaotian/LLaVA-CC3M-Pretrain-595K, and HuggingFaceM4/FairFace. The labels are in Turkish and the dataset is in an instruction-tuning format with separate columns for prompts and completion labels. The original labels (except… See the full description on the dataset page: https://huggingface.co/datasets/ucsahin/Turkish-VLM-Mix-Benchmark.imageimage-to-text10K<n<100K7 likes110 downloads2y agoHugging Face09macular /diabetic-retinopathy-screening-benchmark-africa DR-Africa-Benchmark — Screening-Prevalence-Corrected, Fairness-Instrumented DR Evaluation An evaluation benchmark for diabetic-retinopathy grading under African screening conditions. It does not introduce new labels; it introduces evaluation validity — per-record importance weights that reweight a referral-skewed image set to real Sub-Saharan-Africa population prevalence, plus synthetic subgroup metadata for fairness reporting. Version 1.0.0 · core dr_synth 1.0.0 · part of the… See the full description on the dataset page: https://huggingface.co/datasets/macular/diabetic-retinopathy-screening-benchmark-africa.imageimage-classification1K<n<10K0 likes69 downloads4mo agoHugging Face10nutrientdocs /doc-split-benchmark Doc-Split Benchmark The evaluation slice for page-stream segmentation — the exact set behind the leaderboard and the cloud-VLM comparison. Self-contained (page images embedded), with a reference scorer so results are reproducible. This is the benchmark, not the training corpus (which stays private). 🏆 Leaderboard: doc-split-leaderboard 🎯 Demo: doc-split-demo 🟢 Model: doc-split-mini-e5 (open weights) 🌍 OpenPSS cuts: openpss-mirror (SHORT/LONG, self-contained)… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/doc-split-benchmark.imageimage-classificationn<1K0 likes57 downloads1mo agoHugging Face11DebdipCS /Latent-Resonance-AI-Image-Forensics-Benchmark-N100 Latent Resonance: SOTA Empirical AI Image Forensics Benchmark (N=100 & N=1,000 Scale) Author: Debdip Bandyopadhyay (Independent AI Researcher, Kolkata, India; M.Tech, IIT Jodhpur, AI & Data Science)Preprint & Paper: Latent Resonance: Zero-Shot Autoencoder Inversion and Azimuthal Spectral Forensics for Diffusion Image Attribution (IEEE Flagship / CERN Zenodo 2026) Benchmark Overview This repository provides: The official verified $N=100$ ground-truth image… See the full description on the dataset page: https://huggingface.co/datasets/DebdipCS/Latent-Resonance-AI-Image-Forensics-Benchmark-N100.imageimage-classificationn<1K0 likes53 downloads10d agoHugging Face12Robo531 /ai-detector-benchmark-test-data 🎯 AI Detector Benchmark Test Dataset A comprehensive benchmark dataset for testing AI image detection models. 📊 Dataset Summary Total Images: 700 AI-Generated: 250 images (from 5 different generators) Real Images: 450 images (from 9 diverse datasets) Perfect for: ✅ Testing AI detection models ✅ Creating leaderboards ✅ Comparing model performance ✅ Benchmarking new approaches 🤖 AI Generators Included Generator Images Accuracy Baseline FLUX… See the full description on the dataset page: https://huggingface.co/datasets/Robo531/ai-detector-benchmark-test-data.imageimage-classificationn<1K0 likes52 downloads6mo agoHugging Face13nutrientdocs /document-classification-benchmark Document Classification Benchmark (open-vocab, zero-shot) Given a document image and an arbitrary set of text labels, which one is right? A held-out, zero-shot, open-vocabulary evaluation for document-type classification — labels are supplied at inference, not baked into a head. Test split only; not for training. Every image is drawn from a permissively-licensed, redistributable source. Powers the document-classification-leaderboard and evaluates document-classification-v2… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/document-classification-benchmark.imagezero-shot-image-classification1K<n<10K0 likes51 downloads1mo agoHugging Face14Rapidata /Face_Generation_Benchmark Rapidata Human Face Generation Alignment This T2I dataset contains over ~22'000 human responses, collected in less than 1h using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation. Evaluating 12 different image generation models on which one can generate faces more accurately. The question that the annotators get asked is: "Which Image follows the description of the human better?" To evaluate your own models and create leaderboard check out our… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Face_Generation_Benchmark.imagetext-to-image1K<n<10K16 likes38 downloads11mo agoHugging Face15EDAnonSubmission /benchmark EditJudge-Bench EditJudge-Bench is a synthetic benchmark for auditing vision-language models used as automated judges for image-edit verification. Each row contains a source image, an edited image, a factual edit instruction, counterfactual instructions, and ground-truth scene parameters produced by a controlled Blender/Infinigen generation pipeline. This repository is an anonymous review release for a NeurIPS Evaluations and Datasets submission. Dataset Contents 1… See the full description on the dataset page: https://huggingface.co/datasets/EDAnonSubmission/benchmark.imageimage-classification1K<n<10K0 likes38 downloads5mo agoHugging Face16nutrientdocs /doc-openvocab-benchmark Open-Vocab Document & Figure Classification Benchmark Given a document or figure image and an arbitrary set of text labels, which one is right? This is a zero-shot, open-vocabulary image-classification benchmark for the document-AI setting: every image is scored against a broad ~48-label candidate vocabulary (document types + figure/zone types), and the task is to pick the correct label. The labels are supplied at inference — which is precisely what a fixed-label supervised… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/doc-openvocab-benchmark.imagezero-shot-image-classification1K<n<10K1 likes34 downloads2mo agoHugging Face17ashleyscruse /noble-ai-evidence-benchmark NOBLE AI-Generated Evidence Detection Benchmark A domain-specific benchmark for evaluating AI-generated image detection tools on law enforcement imagery (surveillance footage, bodycam, evidence-style photos). The benchmark spans multiple generator architectures and three image-quality levels designed to mimic the conditions in which real evidence reaches courtrooms. Status: v1.1 release. All three generators (FLUX-schnell, Realistic Vision 5.1, SDXL) complete, paired-prompt design… See the full description on the dataset page: https://huggingface.co/datasets/ashleyscruse/noble-ai-evidence-benchmark.imageimage-classification1K<n<10K0 likes33 downloads4mo agoHugging Face18ash12321 /ai-detector-benchmark-test-data 🎯 AI Detector Benchmark Test Dataset A comprehensive benchmark dataset for testing AI image detection models. 📊 Dataset Summary Total Images: 700 AI-Generated: 250 images (from 5 different generators) Real Images: 450 images (from 9 diverse datasets) Perfect for: ✅ Testing AI detection models ✅ Creating leaderboards ✅ Comparing model performance ✅ Benchmarking new approaches 🤖 AI Generators Included Generator Images Accuracy Baseline FLUX… See the full description on the dataset page: https://huggingface.co/datasets/ash12321/ai-detector-benchmark-test-data.imageimage-classificationn<1K0 likes28 downloads9mo agoHugging Face19BDRC /tibetan-script-classification-benchmark Tibetan Script Classification Benchmark Holdout benchmark for 6-class Tibetan script classification. Test split only — not used during training. All images are BDRC manuscript page scans, balanced by subclass. Class Images Subclasses Danyig 60 DraDring: 25, DraRing: 9, Drathung: 17, Gongshabma: 3, Tsegdrig: 6 Druma 60 Dhumri: 22, DruDring: 20, DruRing: 10, Druchen: 2, Druthung: 6 Gyuyig 60 Khyuyig: 31, Tsumachug: 15, Yigchung: 14 Pedri 60 Peri: 44, Petsuk: 16… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-script-classification-benchmark.imageimage-classificationn<1K0 likes28 downloads3mo agoHugging Face20sukiewang /poi-benchmark POI Benchmark: Multi-City Multimodal Points of Interest A large-scale multimodal benchmark pairing Points of Interest (POIs) with street-view imagery, aerial grid photos, and satellite imagery across 10 major cities on 3 continents. Cities Beijing, Chengdu, Guangzhou, Hong Kong, Shanghai, Shenzhen, London, Melbourne, New York, Sydney. Contents Path Size Type Description metadata_aligned.tar 8.6 GB 11 JSON files Enriched & aligned POI metadata per… See the full description on the dataset page: https://huggingface.co/datasets/sukiewang/poi-benchmark.imageimage-to-text100K<n<1M0 likes27 downloads5mo agoHugging Face21oliveirabruno01 /sheep-facial-expression-benchmark Sheep Facial Expression Benchmark Prepared OpenFARM sheep facial-expression benchmark data from the public Mendeley Data record 10.17632/y5sm4smnfr.5. Source Source dataset: https://data.mendeley.com/datasets/y5sm4smnfr Source DOI: 10.17632/y5sm4smnfr.5 Related paper DOI: 10.1016/j.compag.2020.105528 License: CC BY 4.0 Splits { "train": 172, "test": 74, "train_raw": 898, "test_raw": 225 } train and test are filtered/balanced views for benchmark and… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/sheep-facial-expression-benchmark.imageimage-classification1K<n<10K0 likes20 downloads4mo agoHugging Face22dcher95 /multi-species-benchmark multi-species benchmark Photographs where 2+ species appear in the same frame. Designed to evaluate multi-label species identification and steering capabilities of biological vision-language models. Two sources, unified into one parquet schema. Sources inat21_multilabel (299 rows, 147 images) In-distribution: drawn from iNat21 validation images that already carry an iNat-supplied primary label. We use InternVL3-AWQ to surface images that also… See the full description on the dataset page: https://huggingface.co/datasets/dcher95/multi-species-benchmark.imageimage-classification1K<n<10K0 likes9 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.