MapEval/MapEval-Visual
MapEval-Visual This dataset was introduced in MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models Example Query I am presently visiting Mount Royal Park . Could you please inform me about the nearby historical landmark? Options Circle Stone Secret pool Maison William Caldwell Cottingham Poste de cavalerie du Service de police de la Ville de Montreal Correct Option Circle Stone… See the full description on the dataset page: https://huggingface.co/datasets/MapEval/MapEval-Visual.
2514
1---2license: apache-2.03task_categories:4- multiple-choice5- visual-question-answering6language:7- en8size_categories:9- n<1K10configs:11- config_name: benchmark12 data_files:13 - split: test14 path: dataset.json15paperswithcode_id: mapeval-visual16tags:17- geospatial18---19 20# MapEval-Visual21 22This dataset was introduced in [MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models](https://arxiv.org/abs/2501.00316)23 24# Example25 2627 28#### Query29I am presently visiting Mount Royal Park . Could you please inform me about the nearby historical landmark?30 31#### Options321. Circle Stone332. Secret pool343. Maison William Caldwell Cottingham354. Poste de cavalerie du Service de police de la Ville de Montreal36 37#### Correct Option381. Circle Stone39 40# Prerequisite41 42Download the [Vdata.zip](https://huggingface.co/datasets/MapEval/MapEval-Visual/resolve/main/Vdata.zip?download=true) and extract in the working directory. This directory contains all the images.43 44# Usage45```python46from datasets import load_dataset47import PIL.Image48# Load dataset49ds = load_dataset("MapEval/MapEval-Visual", name="benchmark")50 51for item in ds["test"]:52 53 # Start with a clear task description54 prompt = (55 "You are a highly intelligent assistant. "56 "Based on the given image, answer the multiple-choice question by selecting the correct option.\n\n"57 "Question:\n" + item["question"] + "\n\n"58 "Options:\n"59 )60 61 # List the options more clearly62 for i, option in enumerate(item["options"], start=1):63 prompt += f"{i}. {option}\n"64 65 # Add a concluding sentence to encourage selection of the answer66 prompt += "\nSelect the best option by choosing its number."67 68 # Load image from Vdata/ directory69 img = PIL.Image.open(item["context"])70 71 # Use the prompt as needed72 print([prompt, img]) # Replace with your processing logic73 74 # Then match the output with item["answer"] or item["options"][item["answer"]-1]75 # If item["answer"] == 0: then it's unanswerable76```77 78# Leaderboard79 80| Model | Overall | Place Info | Nearby | Routing | Counting | Unanswerable |81|---------------------------|:-------:|:----------:|:------:|:-------:|:--------:|:------------:|82| Claude-3.5-Sonnet | **61.65** | **82.64** | 55.56 | **45.00** | **47.73** | **90.00** |83| GPT-4o | 58.90 | 76.86 | **57.78** | 50.00 | **47.73** | 40.00 |84| Gemini-1.5-Pro | 56.14 | 76.86 | 56.67 | 43.75 | 32.95 | 80.00 |85| GPT-4-Turbo | 55.89 | 75.21 | 56.67 | 42.50 | 44.32 | 40.00 |86| Gemini-1.5-Flash | 51.94 | 70.25 | 56.47 | 38.36 | 32.95 | 55.00 |87| GPT-4o-mini | 50.13 | 77.69 | 47.78 | 41.25 | 28.41 | 25.00 |88| Qwen2-VL-7B-Instruct | 51.63 | 71.07 | 48.89 | 40.00 | 40.91 | 40.00 |89| Glm-4v-9b | 48.12 | 73.55 | 42.22 | 41.25 | 34.09 | 10.00 |90| InternLm-Xcomposer2 | 43.11 | 70.41 | 48.89 | 43.75 | 34.09 | 10.00 |91| MiniCPM-Llama3-V-2.5 | 40.60 | 60.33 | 32.22 | 32.50 | 31.82 | 30.00 |92| Llama-3-VILA1.5-8B | 32.99 | 46.90 | 32.22 | 28.75 | 26.14 | 5.00 |93| DocOwl1.5 | 31.08 | 43.80 | 23.33 | 32.50 | 27.27 | 0.00 |94| Llava-v1.6-Mistral-7B-hf | 31.33 | 42.15 | 28.89 | 32.50 | 21.59 | 15.00 |95| Paligemma-3B-mix-224 | 30.58 | 37.19 | 25.56 | 38.75 | 23.86 | 10.00 |96| Llava-1.5-7B-hf | 20.05 | 22.31 | 18.89 | 13.75 | 28.41 | 0.00 |97| Human | 82.23 | 81.67 | 82.42 | 85.18 | 78.41 | 65.00 |98 99# Citation100 101If you use this dataset, please cite the original paper:102 103```104@article{dihan2024mapeval,105 title={MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models},106 author={Dihan, Mahir Labib and Hassan, Md Tanvir and Parvez, Md Tanvir and Hasan, Md Hasebul and Alam, Md Almash and Cheema, Muhammad Aamir and Ali, Mohammed Eunus and Parvez, Md Rizwan},107 journal={arXiv preprint arXiv:2501.00316},108 year={2024}109}110```