CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01thoughtworks /document-processing-benchmark Document Processing Benchmark 8 public document datasets (receipts, invoices, forms, bank statements, multi-page docs, contracts) normalized into one parquet schema. Each row has the document, ground-truth annotations, and per-row token/latency/cost numbers from real API calls to one or more reference models. You can read off a target's cost/latency/quality without re-running it. from datasets import load_dataset ds = load_dataset("thoughtworks/document-processing-benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/document-processing-benchmark.tabularimage-to-text10K<n<100K1 likes60 downloads4mo agoHugging Face02DebdipCS /Latent-Resonance-AI-Image-Forensics-Benchmark-N1000 Latent Resonance: SOTA Large-Scale AI Image Forensics Benchmark (N=1,000) Author: Debdip Bandyopadhyay (Independent AI Researcher, Kolkata, India; M.Tech, IIT Jodhpur, AI & Data Science)Preprint & Paper: Latent Resonance: Zero-Shot Autoencoder Inversion and Azimuthal Spectral Forensics for Diffusion Image Attribution (IEEE Flagship / CERN Zenodo 2026) 1. Executive Summary & Diagnostic Suite This repository contains the complete empirical evaluation records… See the full description on the dataset page: https://huggingface.co/datasets/DebdipCS/Latent-Resonance-AI-Image-Forensics-Benchmark-N1000.tabularimage-classification1K<n<10K0 likes49 downloads10d agoHugging Face03EDAnonSubmission /benchmark EditJudge-Bench EditJudge-Bench is a synthetic benchmark for auditing vision-language models used as automated judges for image-edit verification. Each row contains a source image, an edited image, a factual edit instruction, counterfactual instructions, and ground-truth scene parameters produced by a controlled Blender/Infinigen generation pipeline. This repository is an anonymous review release for a NeurIPS Evaluations and Datasets submission. Dataset Contents 1… See the full description on the dataset page: https://huggingface.co/datasets/EDAnonSubmission/benchmark.imageimage-classification1K<n<10K0 likes38 downloads5mo agoHugging Face04dcher95 /multi-species-benchmark multi-species benchmark Photographs where 2+ species appear in the same frame. Designed to evaluate multi-label species identification and steering capabilities of biological vision-language models. Two sources, unified into one parquet schema. Sources inat21_multilabel (299 rows, 147 images) In-distribution: drawn from iNat21 validation images that already carry an iNat-supplied primary label. We use InternVL3-AWQ to surface images that also… See the full description on the dataset page: https://huggingface.co/datasets/dcher95/multi-species-benchmark.imageimage-classification1K<n<10K0 likes9 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.