datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WEIRD
WEIRD
Описание задачи
WEIRD – это расширенная версия подзадачи бинарной классификации оригинального английского бенчмарка WHOOPS!. Датасет оценивает, способна ли мультимодальная модель обнаруживать нарушения здравого смысла в изображениях. Здесь нарушение здравого смысла – это ситуации, противоречащие типичным нормам реальности. Например, пингвины не могут летать, дети не водят автомобили, посетители не накладывают еду официантам, и так далее. В датасете поровну… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/WEIRD.nanopath-evals
NanoPath evaluation data
This is the immutable data mirror used by NanoPath probe protocol v2. It contains only the exact development records consumed by medarc/nanopath: selected THUNDER training/validation images, prepared development-only slide caches, and the two PathoROB subsets. manifest.json records SHA-256 checksums and binds the snapshot to the checked-in benchmark manifests.
No official THUNDER, HEST, or CPTAC classification test record is included. HEST is absent.… See the full description on the dataset page: https://huggingface.co/datasets/medarc/nanopath-evals.tiktok-techjam-2026-eval
TikTok TechJam 2026 Eval
Held-out demonstration pair used by Seer:
COCO val2017 photographs plus the WildFake DALL·E Advanced (DALL·E 3) subset.
Do not train on this split.
Contents
label
meaning
count
origin
0 / real
photograph
5,000
COCO val2017
1 / fake
AI-generated
8,843
WildFake DALL·E Advanced
Columns: image, label, source, generator, id.
id is the COCO stem for reals, and {session}_{stem} for fakes so duplicate
WildFake basenames stay… See the full description on the dataset page: https://huggingface.co/datasets/glennwuwu/tiktok-techjam-2026-eval.trace-rx-eval-predictions
TRACE-RX Evaluation Predictions
Per-image detector scores from an independent evaluation of the two TechJam 2026 TRACE-RX
detectors, run 30 Aug – 1 Sep 2026.
No images here. Every file contains scores, labels, asset ids and transform names only — this is
derived evaluation metadata, not a redistribution of any source imagery. The underlying corpora
(Joshyxwa/data_draft, Joshyxwa/techjam2026, techjam-aigc/wildfake-eval-subset) keep their own
terms, and data_draft's WildFake rows… See the full description on the dataset page: https://huggingface.co/datasets/joelleoqiyi/trace-rx-eval-predictions.lipika-eval
Lipika eval — Indic font recognition benchmark
The frozen validation set behind loopdesk-ai/lipika
(Indic font recognizer): 6,876 synthetic text crops covering 553 freely-licensed font
families across 13 scripts (Devanagari, Bengali, Gujarati, Gurmukhi, Kannada, Malayalam,
Meetei Mayek, Odia, Ol Chiki, Perso-Arabic, Tamil, Telugu, Latin).
This is the set reported as "synthetic val" in the model card (Lipika v2.4 scores 0.849
family top-1 / 0.977 top-5 / 0.991 script). Use it to… See the full description on the dataset page: https://huggingface.co/datasets/loopdesk-ai/lipika-eval.wonders-of-world-images-hf
🌍 Wonders of the World Images 🏛️
¡Bienvenido/a a un viaje visual por las maravillas del mundo!
Este dataset contiene imágenes de 12 maravillas icónicas, listas para que entrenes modelos de visión por computadora, juegues a ser explorador o simplemente disfrutes de la diversidad arquitectónica y natural del planeta.
📦 Estructura del dataset
Clases:
Burj Khalifa
Chichen Itza
Christ the Redeemer
Eiffel Tower
Great Wall of China
Machu Picchu
Pyramids of Giza
Roman… See the full description on the dataset page: https://huggingface.co/datasets/evalverden/wonders-of-world-images-hf.piyoshogi-eval
PiyoShogi Eval (paired, 4 devices)
ぴよ将棋の盤面認識モデルの評価用データセット。4機種の実機スクリーンショットを SFEN 単位で束ねた横持ち形式。
対応機種
機種名
識別子 (devices 列)
画面解像度
iPhone 8
iPhone10,1
750 × 1334
iPhone XR
iPhone11,8
828 × 1792
iPhone 15
iPhone15,4
1179 × 2556
iPad Air M3
iPad14,10
1640 × 2360
paired config
1 行 = 1 SFEN、4 機種分の画像を list で持つ。
Column
Type
説明
sfen
string
SFEN形式の局面文字列
hash
string
SFENのSHA-256
type
string
局面ソース種別(現状は全て existing = ぴよ将棋プリセット由来)… See the full description on the dataset page: https://huggingface.co/datasets/ultemica/piyoshogi-eval.facepass_eval
FacePass Evaluation Dataset (Real LFW Faces)
This dataset contains real face images from the LFW (Labeled Faces in the Wild) dataset, curated for face recognition evaluation.
⚠️ IMPORTANT: This is the corrected version with actual face photographs (not colored squares).
Key Features
✅ Real faces: Actual photographs of people, not synthetic images✅ Balanced dataset: All individuals have 20+ images✅ Proper splits: 80/20 train/test split per person✅ Standardized: Resized to… See the full description on the dataset page: https://huggingface.co/datasets/besartshyti/facepass_eval.evals-eastrus-vl
evals-eastrus-vl
Independent evaluation dataset for the EstrusVision cattle estrus detection model. Contains ground-truth labels for measuring deployment readiness.
Contents
43 total samples (40 in-domain cattle vulval images, 3 out-of-domain)
Embedded image column (no external file dependencies)
Six symptom ground-truth labels per in-domain sample
out_of_domain flag for rejection testing
notes field with clinical observations
Label distribution (in-domain only)… See the full description on the dataset page: https://huggingface.co/datasets/prapaa/evals-eastrus-vl.evaluation
Skill-Aligned Annotation for Text-to-Image Evaluation
Companion dataset for the NeurIPS 2026 paper "Towards Objective Evaluation".
The dataset contains generated images from 7 text-to-image models, evaluated
by 6 human annotators (anonymized) plus an LLM judge across 9 skill-aligned
annotation strategies.
Configs
Config
Rows
Description
images
621
Generated images (621 WebP) with embedded bytes; one row per (prompt_id, generator).
prompts
179
Per-prompt… See the full description on the dataset page: https://huggingface.co/datasets/Skill-Aigned/evaluation.
