CoolFace
Datasetpublic

EDAnonSubmission/benchmark

EditJudge-Bench EditJudge-Bench is a synthetic benchmark for auditing vision-language models used as automated judges for image-edit verification. Each row contains a source image, an edited image, a factual edit instruction, counterfactual instructions, and ground-truth scene parameters produced by a controlled Blender/Infinigen generation pipeline. This repository is an anonymous review release for a NeurIPS Evaluations and Datasets submission. Dataset Contents… See the full description on the dataset page: https://huggingface.co/datasets/EDAnonSubmission/benchmark.

sourceHugging Facecc-by-nc-4.0updated 5mo agoView on Hugging Face
0likes42downloads
1 commits on main
7f5115e5mo ago

Add EditJudge-Bench dataset

anon