CoolFace
15 results

mmbench

lmms-lab-encoder /MMBenchimage10K<n<100K25 likes19k downloads3y agoHugging Facedeepvk /MMBench-ru MMBench-ru This is a translated version of original MMBench dataset and stored in format supported for lmms-eval pipeline. For this dataset, we: Translate the original one with gpt-4o Filter out unsuccessful translations, i.e. where the model protection was triggered Manually validate most common errors Dataset Structure Dataset includes only dev split that is translated from dev split in lmms-lab/MMBench_EN. Dataset contains 3910 samples in the same to… See the full description on the dataset page: https://huggingface.co/datasets/deepvk/MMBench-ru.imagevisual-question-answering1K<n<10K6 likes2.6k downloads2y agoHugging Facenicklashansen /mmbench2 MMBench2 Hallucination in World Models is Predictable and Preventable Nicklas Hansen &nbsp;·&nbsp; Xiaolong Wang &nbsp;·&nbsp; UC San Diego MMBench2 is a large-scale dataset for visual world modeling, accompanying the paper Hallucination in World Models is Predictable and Preventable. It spans 210 continuous control tasks across 10 domains (DMControl, DMControl Extended, Meta-World, ManiSkill3, MuJoCo, MiniArcade, Box2D, RoboDesk, OGBench, and Atari) comprising 65,600… See the full description on the dataset page: https://huggingface.co/datasets/nicklashansen/mmbench2.reinforcement-learning10K<n<100K1 likes2.3k downloads3mo agoHugging Facemmbench /MM-SpuBench MM-SpuBench Datacard Basic Information Title: The Multimodal Spurious Benchmark (MM-SpuBench) Description: MM-SpuBench is a comprehensive benchmark designed to evaluate the robustness of MLLMs to spurious biases. This benchmark systematically assesses how well these models distinguish between core and spurious features, providing a detailed framework for understanding and quantifying spurious biases. Data Structure: ├── data/images │ ├── 000000.jpg │ ├── 000001.jpg │… See the full description on the dataset page: https://huggingface.co/datasets/mmbench/MM-SpuBench.imagequestion-answering1K<n<10K2 likes2.3k downloads2y agoHugging Facelmms-lab /MMBench_EN Dataset Card for "MMBench_EN" Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of the English subset of MMBench. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @article{MMBench, author = {Yuan Liu, Haodong… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/MMBench_EN.image10K<n<100K7 likes1.7k downloads3y agoHugging FaceHuggingFaceM4 /MMBench Dataset Card for "MMBench" More Information needed text10K<n<100K5 likes802 downloads2y agoHugging Face