datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
polymath
Paper Information
We present PolyMATH, a challenging benchmark aimed at evaluating the general cognitive reasoning abilities of MLLMs.
PolyMATH comprises 5,000 manually collected high-quality images of cognitive textual and visual challenges across 10 distinct categories, including pattern recognition, spatial reasoning, and relative reasoning.
We conducted a comprehensive, and quantitative evaluation of 15 MLLMs using four diverse prompting strategies, including Chain-of-Thought… See the full description on the dataset page: https://huggingface.co/datasets/him1411/polymath.Polymath
