datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gw-tti-playground-queueMJHQ-30K
MJHQ-30K Benchmark
Model
Overall FID
SDXL-1-0-refiner
9.55
playground-v2-1024px-aesthetic
7.07
We introduce a new benchmark, MJHQ-30K, for automatic evaluation of a model’s aesthetic quality. The benchmark computes FID on a high-quality dataset to gauge aesthetic quality.
We curate the high-quality dataset from Midjourney with 10 common categories, each category with 3K samples. Following common practice, we use aesthetic score and CLIP score to ensure high image… See the full description on the dataset page: https://huggingface.co/datasets/playgroundai/MJHQ-30K.FSCM_Flood_playgroundcrello-capVisR-Bench
Testing Code:
The GitHub repo for testing code: VisR-Bench
Data Download
This is the page images of all documents of the VisR-Bench dataset.
git lfs install
git clone https://huggingface.co/datasets/puar-playground/VisR-Bench
The code above will download the VisR-Bench folder, which is required for testing.
Reference
@misc{chen2025visrbenchempiricalstudyvisual,
title={VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for… See the full description on the dataset page: https://huggingface.co/datasets/puar-playground/VisR-Bench.CapsBench
CapsBench
CapsBench is a captioning evaluation dataset designed to comprehensively assess the quality of the captions across 17 categories: general,
image type, text, color, position, relation, relative position, entity, entity size, entity shape, count, emotion, blur, image artifacts,
proper noun (world knowledge), color palette, and color grading.
There are 200 images and 2471 questions for them, resulting in 12 questions per image on average. Images represent a wide variety of… See the full description on the dataset page: https://huggingface.co/datasets/playgroundai/CapsBench.MMR
Evaluating Reading Ability of Large Multimodal Models
This is the evaluation data of the MRR Benchmark for Vision Language models.
The bounding boxes are in the format of [x_min, y_min, x_max, y_max] and normalized to ralative ratio between 0 and 1.
Reference
@inproceedings{mmr2024chen,
title={Evaluating Reading Ability of Large Multimodal Models},
author={Chen, Jian and Zhang, Ruiyi and Zhou, Yufan and Gu, Jiuxiang and Rossi, Ryan and Chen, Changyou}… See the full description on the dataset page: https://huggingface.co/datasets/puar-playground/MMR.blip_clipseg_inpainting_ip2p_data_test
Dataset Card for "blip_clipseg_inpainting_ip2p_data_test"
More Information needed
playground2aas_benchmark-playgroundfake_playground-25_1kmm_playgrounddigital-playgroundplayground-v2.5-1024px-aesthetic_woman_class_imagesplayground-v2.5-1024px-aesthetic_man_class_imagesFSCM_Snow_playgroundPlaygrounds_2026
