stanford-crfm/image2struct-musicsheet-v1
Image2Struct - Music Sheet Paper | Website | Datasets (Webpages, Latex, Music sheets) | Leaderboard | HELM repo | Image2Struct repo License: Apache License Version 2.0, January 2004 Dataset description Image2struct is a benchmark for evaluating vision-language models in practical tasks of extracting structured information from images. This subdataset focuses on Music sheets. The model is given an image of the expected output with the prompt: Please generate the… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/image2struct-musicsheet-v1.
Image2Struct - Music Sheet
Paper | Website | Datasets (Webpages, Latex, Music sheets) | Leaderboard | HELM repo | Image2Struct repo
License: Apache License Version 2.0, January 2004
Dataset description
Image2struct is a benchmark for evaluating vision-language models in practical tasks of extracting structured information from images. This subdataset focuses on Music sheets. The model is given an image of the expected output with the prompt:
Please generate the Lilypond code to generate a music sheet that looks like this image as much as feasibly possible.
This music sheet was created by me, and I would like to recreate it using Lilypond.The data was collected from IMSLP and has no ground truth. This means that while we prompt models to output some Lilypond code to recreate the image of the music sheet, we do not have access to a Lilypond code that could reproduce the image and would act as a "ground-truth".
There is no wild subset as this already constitutes a dataset without ground-truths.
Uses
To load the subset music of the dataset to be sent to the model under evaluation in Python:
import datasets
datasets.load_dataset("stanford-crfm/i2s-musicsheet", "music", split="validation")To evaluate a model on Image2Musicsheet (equation) using HELM, run the following command-line commands:
pip install crfm-helm
helm-run --run-entries image2musicsheet,model=vlm --models-to-run google/gemini-pro-vision --suite my-suite-i2s --max-eval-instances 10You can also run the evaluation for only a specific difficulty:
helm-run --run-entries image2musicsheet:difficulty=hard,model=vlm --models-to-run google/gemini-pro-vision --suite my-suite-i2s --max-eval-instances 10For more information on running Image2Struct using HELM, refer to the HELM documentation and the article on reproducing leaderboards.
Citation
BibTeX:
@misc{roberts2024image2struct,
title={Image2Struct: A Benchmark for Evaluating Vision-Language Models in Extracting Structured Information from Images},
author={Josselin Somerville Roberts and Tony Lee and Chi Heem Wong and Michihiro Yasunaga and Yifan Mai and Percy Liang},
year={2024},
eprint={TBD},
archivePrefix={arXiv},
primaryClass={TBD}
}