facebook/Common-O
Common-O measuring multimodal reasoning across scenes Common-O, inspired by cognitive tests for humans, probes multimodal LLMs' ability to reason across scenes by asking "what’s in common?" Common-O is comprised of household objects: We have two subsets: Common-O (3 - 8 objects) and Common-O Complex (8 - 16 objects). Multimodal LLMs excel at single image perception, but struggle with multi-scene reasoning Evaluating a Multimodal LLM on Common-O… See the full description on the dataset page: https://huggingface.co/datasets/facebook/Common-O.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face