CoolFace
Datasetpublic

Eventual-Inc/VLADBench-reeval

VLADBench-reeval We re-evaluated the VLADBench benchmark (Li et al., 2025, arXiv:2503.21505) against current SOTA VLMs under the original scoring criteria and prompts. See Eventual-Inc/VLADBench for the code, the companion site for interactive results, and the article for a summary of the findings. Cost vs Score Up and to the left is better. The dashed line is the cost-performance frontier which is defined by measuring the best TOTAL score available against its… See the full description on the dataset page: https://huggingface.co/datasets/Eventual-Inc/VLADBench-reeval.

sourceHugging Faceotherupdated 2d agoView on Hugging Face
2likes257downloads
13 commits on main
5ae1ae12d ago

Results as of 2026-09-23T04:13:26Z: 15 models, 167,895 answers

EverettKleven
f56797d2d ago

Card: link the article; full author list in the citation

EverettKleven
53699884d ago

Results as of 2026-09-23T04:13:26Z: 15 models, 167,895 answers

EverettKleven
be94b154d ago

Results as of 2026-09-22T20:31:36Z: 13 models, 145,509 answers

EverettKleven
3c93bb34d ago

Results as of 2026-09-22T20:31:36Z: 13 models, 145,509 answers

EverettKleven
a0591ee10d ago

Results as of 2026-09-16T04:34:41Z: 12 models, 134,316 answers

EverettKleven
cb301a610d ago

Results as of 2026-09-16T04:34:41Z: 12 models, 134,316 answers

EverettKleven
2105e5210d ago

Results as of 2026-09-16T04:34:41Z: 12 models, 134,316 answers

EverettKleven
660b74e10d ago

Results as of 2026-09-16T04:34:41Z: 12 models, 134,316 answers

EverettKleven
57251e910d ago

Results as of 2026-09-16T04:34:41Z: 12 models, 134,316 answers

EverettKleven
4e9816410d ago

Results as of 2026-09-16T04:34:41Z: 12 models, 134,316 answers

EverettKleven
f1a05a110d ago

Results as of 2026-09-16T04:34:41Z: 12 models, 134,316 answers

EverettKleven
1d0328710d ago

initial commit

EverettKleven