MaCBench
Datasets
All datasets matching “MaCBench”MaCBench
MaCBench
A Chemistry and Materials Benchmark for evaluating Vision Large Language Models
⚠️ IMPORTANT NOTICE - NOT FOR TRAINING
🚫 THIS DATASET IS STRICTLY FOR EVALUATION PURPOSES ONLY 🚫
DO NOT USE THIS DATASET FOR TRAINING OR FINE-TUNING MODELS
This benchmark is designed exclusively for evaluation and testing of existing models. Using this data for training would compromise the integrity of the benchmark and invalidate evaluation results. Please… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/MaCBench.MaCBench-Ablations
MaCBench-Ablations
A Chemistry and Materials Benchmark for evaluating Vision Large Language Models
⚠️ IMPORTANT NOTICE - NOT FOR TRAINING
🚫 THIS DATASET IS STRICTLY FOR EVALUATION PURPOSES ONLY 🚫
DO NOT USE THIS DATASET FOR TRAINING OR FINE-TUNING MODELS
This benchmark is designed exclusively for evaluation and testing of existing models. Using this data for training would compromise the integrity of the benchmark and invalidate evaluation… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/MaCBench-Ablations.MAC_Bench
MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding
📋 Dataset Description
MAC is a comprehensive live benchmark designed to evaluate multimodal large language models (MLLMs) on scientific understanding tasks. The dataset focuses on scientific journal cover understanding, providing challenging testbeds for assessing visual-textual comprehension capabilities of MLLMs in academic domains.
🎯 Tasks
1. Image-to-Text… See the full description on the dataset page: https://huggingface.co/datasets/mhjiang0408/MAC_Bench.MaCBench-ResultsMaCBench-Prompt-Ablations
MaCBench-Prompt-Ablations
A Chemistry and Materials Benchmark for evaluating Vision Large Language Models
⚠️ IMPORTANT NOTICE - NOT FOR TRAINING
🚫 THIS DATASET IS STRICTLY FOR EVALUATION PURPOSES ONLY 🚫
DO NOT USE THIS DATASET FOR TRAINING OR FINE-TUNING MODELS
This benchmark is designed exclusively for evaluation and testing of existing models. Using this data for training would compromise the integrity of the benchmark and invalidate… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/MaCBench-Prompt-Ablations.macbench-dev
