CoolFace
Datasetpublic

stemdataset/STEM

STEM Dataset ๐Ÿ“ƒ [Paper] โ€ข ๐Ÿ’ป [Github] โ€ข ๐Ÿค— [Dataset] โ€ข ๐Ÿ† [Leaderboard] โ€ข ๐Ÿ“ฝ [Slides] โ€ข ๐Ÿ“‹ [Poster] This dataset is proposed in the ICLR 2024 paper: Measuring Vision-Language STEM Skills of Neural Models. We introduce a new challenge to test the STEM skills of neural models. The problems in the real world often require solutions, combining knowledge from STEM (science, technology, engineering, and math). Unlike existing datasets, our dataset requires the understanding ofโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/stemdataset/STEM.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
6likes1.3kdownloads
Dataset Card

STEM Dataset

<p align="center"> ๐Ÿ“ƒ <a href="https://arxiv.org/abs/2402.17205" target="blank">[Paper]</a> โ€ข ๐Ÿ’ป <a href="https://github.com/stemdataset/STEM" target="blank">[Github]</a> โ€ข ๐Ÿค— <a href="https://huggingface.co/datasets/stemdataset/STEM" target="blank">[Dataset]</a> โ€ข ๐Ÿ† <a href="https://huggingface.co/spaces/stemdataset/stem-leaderboard" target="blank">[Leaderboard]</a> โ€ข ๐Ÿ“ฝ <a href="https://github.com/stemdataset/STEM/blob/main/assets/STEM-Slides.pdf" target="blank">[Slides]</a> โ€ข ๐Ÿ“‹ <a href="https://github.com/stemdataset/STEM/blob/main/assets/poster.pdf" target="blank">[Poster]</a> </p>

This dataset is proposed in the ICLR 2024 paper: Measuring Vision-Language STEM Skills of Neural Models. We introduce a new challenge to test the STEM skills of neural models. The problems in the real world often require solutions, combining knowledge from STEM (science, technology, engineering, and math). Unlike existing datasets, our dataset requires the understanding of multimodal vision-language information of STEM. Our dataset features one of the largest and most comprehensive datasets for the challenge. It includes 448 skills and 1,073,146 questions spanning all STEM subjects. Compared to existing datasets that often focus on examining expert-level ability, our dataset includes fundamental skills and questions designed based on the K-12 curriculum. We also add state-of-the-art foundation models such as CLIP and GPT-3.5-Turbo to our benchmark. Results show that the recent model advances only help master a very limited number of lower grade-level skills (2.5% in the third grade) in our dataset. In fact, these models are still well below (averaging 54.7%) the performance of elementary students, not to mention near expert-level performance. To understand and increase the performance on our dataset, we teach the models on a training split of our dataset. Even though we observe improved performance, the model performance remains relatively low compared to average elementary students. To solve STEM problems, we will need novel algorithmic innovations from the community.

Authors

Jianhao Shen, Ye Yuan, Srbuhi Mirzoyan, Ming Zhang, Chenguang Wang

Resources

  • โ€”Code: https://github.com/stemdataset/STEM
  • โ€”Paper: https://arxiv.org/abs/2402.17205
  • โ€”Dataset: https://huggingface.co/datasets/stemdataset/STEM
  • โ€”Leaderboard: https://huggingface.co/spaces/stemdataset/stem-leaderboard

Dataset

The dataset consists of multimodal multi-choice questions. The dataset is splitted into train, valid and test sets. The groundtruth answers of the test set are not released and everyone can submit the test predictions to the leaderboard. The basic statistics of the dataset are as follows:

Subject#Skills#QuestionsAvg. #A#Train#Valid#Test
Science82186,7402.8112,12037,34337,277
Technology98,5664.05,1401,7131,713
Engineering618,9812.512,0553,4403,486
Math351858,8592.8515,482171,776171,601
Total4481,073,1462.8644,797214,272214,077

The dataset is in the following format:

python
DatasetDict({
    train: Dataset({
        features: ['subject', 'grade', 'skill', 'pic_choice', 'pic_prob', 'problem', 'problem_pic', 'choices', 'choices_pic', 'answer_idx'],
        num_rows: 644797
    })
    valid: Dataset({
        features: ['subject', 'grade', 'skill', 'pic_choice', 'pic_prob', 'problem', 'problem_pic', 'choices', 'choices_pic', 'answer_idx'],
        num_rows: 214272
    })
    test: Dataset({
        features: ['subject', 'grade', 'skill', 'pic_choice', 'pic_prob', 'problem', 'problem_pic', 'choices', 'choices_pic', 'answer_idx'],
        num_rows: 214077
    })
})

And the detailed descriptions are as follows:

  • โ€”subject: str
  • โ€”The subject of the question, one of science, technology, engineer, math.
  • โ€”grade: str
  • โ€”The grade level information of the question, e.g., grade-1.
  • โ€”skill: str
  • โ€”The skill level information of the question.
  • โ€”pic_choice: bool
  • โ€”Whether the choices are images.
  • โ€”pic_prob: bool
  • โ€”Whether the question has an image.
  • โ€”problem: str
  • โ€”The question description.
  • โ€”problem_pic: bytes
  • โ€”The image of the question.
  • โ€”choices: Optional[List[str]]
  • โ€”The choices of the question. If pic_choice is True, the choices are images and will be saved into choices_pic, and the choices with be set to None.
  • โ€”choices_pic: Optional[List[bytes]]
  • โ€”The choices images. If pic_choice is False, the choices are strings and will be saved into choices, and the choices_pic with be set to None.
  • โ€”answer_idx: int
  • โ€”The index of the correct answer in the choices or choices_pic. If the split is test, the answer_idx is -1.

The bytes can be easily read by the following code:

python
from PIL import Image
def bytes_to_image(img_bytes: bytes) -> Image:
    img = Image.open(io.BytesIO(img_bytes))
    return img

Example Questions

Questions containing images

*Question: What is the domain of this function?*

*Image*: [image]

*Choices: ["{x | x <= -6}", "all real numbers", "{x | x > 3}", "{x | x >= 0}"]*

*Answer: 1*

*Metadata*:

json
{
  "subject": "math",
  "grade": "algebra-1",
  "skill": "domain-and-range-of-absolute-value-functions-graphs",
  "pic_choice": false,
  "pic_prob": true,
  "problem": "What is the domain of this function?",
  "problem_pic": "b'\\x89PNG\\r\\n\\x1a\\n\\x00\\x00\\x00\\rIHDR\\x00\\x00\\x02\\xd8'...",
  "choices": [
    "$\\{x \\mid x \\leq -6\\}$",
    "all real numbers",
    "$\\{x \\mid x > 3\\}$",
    "$\\{x \\mid x \\geq 0\\}$"
  ],
  "choices_pic": null,
  "answer_idx": 1
}

Choices containing images

*Question: The three scatter plots below show the same data set. Choose the scatter plot in which the outlier is highlighted.*

*Choices*: <div style="display: flex; justify-content: space-between;"> <img src="assets/examplechoicepic0.png" style="width: 30%" /> <img src="assets/examplechoicepic1.png" style="width: 30%" /> <img src="assets/examplechoicepic_2.png" style="width: 30%" /> </div>

*Answer: 1*

*Metadata*:

json
{
  "subject": "math",
  "grade": "precalculus",
  "skill": "outliers-in-scatter-plots",
  "pic_choice": true,
  "pic_prob": false,
  "problem": "The three scatter plots below show the same data set. Choose the scatter plot in which the outlier is highlighted.",
  "problem_pic": null,
  "choices": null,
  "choices_pic": [
    "b'\\x89PNG\\r\\n\\x1a\\n\\x00\\x00\\x00\\rIHDR\\x00\\x00\\x01N'...",
    "b'\\x89PNG\\r\\n\\x1a\\n\\x00\\x00\\x00\\rIHDR\\x00\\x00\\x01N'...",
    "b'\\x89PNG\\r\\n\\x1a\\n\\x00\\x00\\x00\\rIHDR\\x00\\x00\\x01N'..."
  ],
  "answer_idx": 1
}

How to Use

Please refer to our code for the usage of evaluation on the dataset.

Citation

bibtex
@inproceedings{shen2024measuring,
  title={Measuring Vision-Language STEM Skills of Neural Models},
  author={Shen, Jianhao and Yuan, Ye and Mirzoyan, Srbuhi and Zhang, Ming and Wang, Chenguang},
  booktitle={ICLR},
  year={2024}
}

Dataset Card Contact

stemdataset@gmail.com