datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GUIGuard-Bench
GUIGuard-Bench (Public Ladder)
GUIGuard-Bench is a cross-platform GUI agent benchmark for studying privacy risks and privacy-preserving execution in multimodal GUI agents.
This public-ladder release contains 121 GUI interaction trajectories (68 Android + 53 PC) for benchmark evaluation, with 26,407 region-level privacy annotations across 2,002 screenshots.
For the anonymous review version of the evaluation toolkit, see GUIGaurd-Bench-CA4F.
Dataset Summary
GUI agents… See the full description on the dataset page: https://huggingface.co/datasets/ShaofantuoshuzhengzhiSha/GUIGuard-Bench.SA-BENCH
SA-BENCH
SA-BENCH is the benchmark dataset released with “Beyond Pixels: Benchmarking and Reward-Based Assessing Framework for Visual Spatial Aesthetics.”
Accepted to CVPRW 2026.
GitHub | CVF Open Access | arXiv | Model
It evaluates the spatial aesthetics of interior images along four dimensions:
distortion
harmony
layout
lighting
SA-BENCH contains 17,768 annotated examples across four spatial-aesthetic dimensions, with image assets and human annotations for training and… See the full description on the dataset page: https://huggingface.co/datasets/gaoyuan-ai/SA-BENCH.HUGO-Bench-Paper-Reproducibility
HUGO-Bench Paper Reproducibility
Supplementary data and reproducibility materials for the paper:
Vision Transformers for Zero-Shot Clustering of Animal Images: A Comparative Benchmarking Study - https://arxiv.org/abs/2602.03894
Hugo Markoff, Stefan Hein Bengtson, Michael Ørsted
Aalborg University, Denmark
Dataset Description
This repository contains complete experimental results, pre-computed embeddings, and execution logs from our comprehensive benchmarking study… See the full description on the dataset page: https://huggingface.co/datasets/AI-EcoNet/HUGO-Bench-Paper-Reproducibility.document-processing-benchmark
Document Processing Benchmark
8 public document datasets (receipts, invoices, forms, bank statements,
multi-page docs, contracts) normalized into one parquet schema. Each row
has the document, ground-truth annotations, and per-row token/latency/cost
numbers from real API calls to one or more reference models. You can
read off a target's cost/latency/quality without re-running it.
from datasets import load_dataset
ds = load_dataset("thoughtworks/document-processing-benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/document-processing-benchmark.Latent-Resonance-AI-Image-Forensics-Benchmark-N1000
Latent Resonance: SOTA Large-Scale AI Image Forensics Benchmark (N=1,000)
Author: Debdip Bandyopadhyay (Independent AI Researcher, Kolkata, India; M.Tech, IIT Jodhpur, AI & Data Science)Preprint & Paper: Latent Resonance: Zero-Shot Autoencoder Inversion and Azimuthal Spectral Forensics for Diffusion Image Attribution (IEEE Flagship / CERN Zenodo 2026)
1. Executive Summary & Diagnostic Suite
This repository contains the complete empirical evaluation records… See the full description on the dataset page: https://huggingface.co/datasets/DebdipCS/Latent-Resonance-AI-Image-Forensics-Benchmark-N1000.Weitaikang_bench_data
HeroFrame-Bench
A benchmark for key-frame selection that scores a selected frame directly,
without routing it through a question-answering model and without requiring it
to match a fixed reference set.
204 films, 2,031 frozen comparison chains, 1,970 learned criteria, 30,465
pairwise judgements.
What problem this addresses
A film is almost always encountered first as a single still: a cover, a
thumbnail, a poster. Producing that still from the film is the task… See the full description on the dataset page: https://huggingface.co/datasets/weitaikang/Weitaikang_bench_data.CPRT-Bench
Dataset Card for CPRT-Bench
CPRT-Bench is a benchmark dataset for assessing privacy risk in images, designed to model privacy as a graded and composition-dependent phenomenon.
Dataset Details
Dataset Description
The dataset contains approximately 6.7K images annotated with:
Ordinal severity levels (4 levels of privacy risk)
Continuous risk scores (fine-grained privacy assessment)
All images are sourced from the VISPR (Visual Privacy Advisor). CPRT-Bench… See the full description on the dataset page: https://huggingface.co/datasets/timtsapras23/CPRT-Bench.benchmark
EditJudge-Bench
EditJudge-Bench is a synthetic benchmark for auditing vision-language models used as
automated judges for image-edit verification. Each row contains a source image,
an edited image, a factual edit instruction, counterfactual instructions, and
ground-truth scene parameters produced by a controlled Blender/Infinigen
generation pipeline.
This repository is an anonymous review release for a NeurIPS Evaluations and
Datasets submission.
Dataset Contents
1… See the full description on the dataset page: https://huggingface.co/datasets/EDAnonSubmission/benchmark.SA-BENCH
SA-BENCH
SA-BENCH is the benchmark dataset released with “Beyond Pixels: Benchmarking and Reward-Based Assessing Framework for Visual Spatial Aesthetics.”
It evaluates the spatial aesthetics of interior images along four dimensions:
distortion
harmony
layout
lighting
The dataset contains 17,768 annotated examples across four spatial-aesthetic dimensions, with image assets and human annotations for training and evaluation.
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AliHome3D/SA-BENCH.jump-sample
cp-bg-bench preview — jump
Compact preview of the jump dataset from the cp-bg-bench
benchmark. Stratified subset of the full release; designed so that the
held-out-batch perturbation-recall and cp_measure-prediction evals can
be reproduced end-to-end against this small slice alone.
Cells
913
Wells
69
Perturbations
33
Held-out batch
source_4
Views
crops, crops_density, seg, seg_density
What's in this repo
crops/ # HF dataset… See the full description on the dataset page: https://huggingface.co/datasets/cp-bg-bench-anon/jump-sample.multi-species-benchmark
multi-species benchmark
Photographs where 2+ species appear in the same frame. Designed to evaluate
multi-label species identification and steering capabilities of biological
vision-language models. Two sources, unified into one parquet schema.
Sources
inat21_multilabel (299 rows, 147 images)
In-distribution: drawn from iNat21
validation images that already carry an iNat-supplied primary label. We use
InternVL3-AWQ to surface
images that also… See the full description on the dataset page: https://huggingface.co/datasets/dcher95/multi-species-benchmark.synthetic-face-sdxl-instantid-bench
Synthetic Face Detection Benchmark — SDXL+InstantID
Version: v1.0.0 · Build date: 2026-05-16 · Rows: 26492
Evaluation benchmark for synthetic-face detection under platform-realistic
conditions. Sampled to satisfy the ISO/IEC 19795 floor of 300 samples per
demographic subgroup across a 6×2 (skin tone × gender) cell grid. Not
training data; not licensed for commercial use.
See release.json for build provenance, manifest.csv for per-row
license attestation, LICENSES.csv for the… See the full description on the dataset page: https://huggingface.co/datasets/danb21/synthetic-face-sdxl-instantid-bench.
