rhli/genarena
GenArena A unified evaluation framework for visual generation tasks using VLM-based pairwise comparison and Elo ranking. Abstract The rapid advancement of visual generation models has outpaced traditional evaluation approaches, necessitating the adoption of Vision-Language Models as surrogate judges. In this work, we systematically investigate the reliability of the prevailing absolute pointwise scoring standard, across a wide spectrum of visual generation… See the full description on the dataset page: https://huggingface.co/datasets/rhli/genarena.
0130
