datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ScienceQA
Dataset Card Creation Guide
Dataset Summary
Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Supported Tasks and Leaderboards
Multi-modal Multiple Choice
Languages
English
Dataset Structure
Data Instances
Explore more samples here.
{'image': Image,
'question': 'Which of these states is farthest north?',
'choices': ['West Virginia', 'Louisiana', 'Arizona', 'Oklahoma'],
'answer': 0… See the full description on the dataset page: https://huggingface.co/datasets/derek-thomas/ScienceQA.ScienceQA
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of derek-thomas/ScienceQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@inproceedings{lu2022learn,
title={Learn to Explain: Multimodal Reasoning via Thought… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/ScienceQA.vidore_v3_computer_scienceViDoRe V3 : Computer Science
This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark.
About ViDoRe v3
ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science.ScienceQA-IMG
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted and filtered version of derek-thomas/ScienceQA with only image instances. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@inproceedings{lu2022learn,
title={Learn to Explain:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/ScienceQA-IMG.vidore_v3_computer_science_mteb_format
Vidore3ComputerScienceRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Retrieve associated pages according to questions.
Task category
t2i
Domains
Academic
Reference
https://huggingface.co/blog/QuentinJG/introducing-vidore-v3
Source datasets:
vidore/vidore_v3_computer_science
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science_mteb_format.Science-T2I-Fullset
Science-T2I Fullset
Resources
Website
arXiv: Paper
GitHub: Code
Huggingface: SciScore
Huggingface: Science-T2I-S&C Benchmark
Data
The Science-T2I Fullset comprises a comprehensive collection of data for scientific T2I generation, including both training and test sets with a unified data structure. The test sets are split into 'test-S' and 'test-C,' corresponding to the Science-T2I-S and Science-T2I-C benchmarks, respectively.
Download Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Jialuo21/Science-T2I-Fullset.ScienceOlympiad
ScienceOlympiad: Challenging AI with Olympiad-Level Multimodal Science Problems
Dataset Description
The ScienceOlympiad dataset is a meticulously curated benchmark designed to test the limits of current AI models in scientific reasoning. It comprises elite, competition-level problems in physics and chemistry. Addressing the need for more diverse and realistic challenges, ScienceOlympiad introduces multimodal integration as a key dimension. Unlike purely text-based… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/ScienceOlympiad.ScienceQA
Dataset Card for "ScienceQA"
Dataset Summary
ScienceQA is collected from elementary and high school science curricula, and contains 21,208 multimodal multiple-choice science questions. Out of the questions in ScienceQA, 10,332 (48.7%) have an image context, 10,220 (48.2%) have a text context, and 6,532 (30.8%) have both. Most questions are annotated with grounded lectures (83.9%) and detailed explanations (90.5%). The lecture and explanation provide general external… See the full description on the dataset page: https://huggingface.co/datasets/TheMrguiller/ScienceQA.ScienceQA-LLAVA
Dataset Card for "ScienceQA-LLAVA"
More Information needed
scienceqa-problems-datasetscience-r1DiEm_HTR
Dataset Card for DiEm HTR
The DiEm HTR dataset is a ground truth dataset for historical danish handwriting in the 17th and 18th century, generated as part of the Digitalisering af Enesteministerialbøger project at the Danish National Archives.
Dataset Details
Dataset Description
The Digitalisering af Enesteministerialbøger project (DiEm) at the Danish National Archives aims to transcribe and make publically available all of the danish parish registers from… See the full description on the dataset page: https://huggingface.co/datasets/RA-Data-Science/DiEm_HTR.ScienceQA
Dataset Card Creation Guide
Dataset Summary
Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Supported Tasks and Leaderboards
Multi-modal Multiple Choice
Languages
English
Dataset Structure
Data Instances
Explore more samples here.
{'image': Image,
'question': 'Which of these states is farthest north?',
'choices': ['West Virginia', 'Louisiana', 'Arizona', 'Oklahoma'],
'answer': 0… See the full description on the dataset page: https://huggingface.co/datasets/Gisiyuan/ScienceQA.ScienceQAScienceQAImg_Modif
Dataset Card for "ScienceQAImg_Modif"
This dataset contains the ScienceQA benchmark where only examples with an image are kept, and where we formatted the prompt.
ScienceQA
Dataset Card Creation Guide
Dataset Summary
Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Supported Tasks and Leaderboards
Multi-modal Multiple Choice
Languages
English
Dataset Structure
Data Instances
Explore more samples here.
{'image': Image,
'question': 'Which of these states is farthest north?',
'choices': ['West Virginia', 'Louisiana', 'Arizona', 'Oklahoma'],
'answer': 0… See the full description on the dataset page: https://huggingface.co/datasets/aarav7397/ScienceQA.tubitak-science-olympiad-tr
TUBITAK Science Olympiad Dataset
This dataset contains multiple-choice and open-ended scientific questions sourced from the TUBITAK (The Scientific and Technological Research Council of Turkey) Science Olympiads spanning various years. It is intended to serve as a benchmark for evaluating the advanced analytical, mathematical, and computational reasoning capabilities of Large Language Models (LLMs) in the Turkish language.
The dataset comprises approximately 2700 problems across… See the full description on the dataset page: https://huggingface.co/datasets/ytu-ce-cosmos/tubitak-science-olympiad-tr.modern-danish-handwriting
Dataset Card for Modern Danish Handwriting
The Modern Danish Handwriting dataset is a Danish-language dataset containing more than 200 pages of transcribed and proofread handwritten text.
Dataset Details
Dataset Description
The Modern Danish Handwriting dataset currently consists of handwritten samples of text from the ePAROLE dataset. The samples were created by volunteers at the Danish National Archives and guests at the festival Historiske Dage in 2025.… See the full description on the dataset page: https://huggingface.co/datasets/RA-Data-Science/modern-danish-handwriting.vidore_v3_computer_science_embeddingNOTE
ViDoRe V3: Computer Science dataset ColQwen2 Embeddings
This dataset contains pre-computed embeddings for the ViDoRe V3 : Computer Science dataset using the ColQwen2 model.
ViDoRe V3 : Computer Science
This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark.
About ViDoRe v3
ViDoRe V3 is our latest benchmark for RAG evaluation on… See the full description on the dataset page: https://huggingface.co/datasets/WenxingZhu/vidore_v3_computer_science_embedding.Science-T2I
Science-T2I Benchmark
Resources
Website
arXiv: Paper
GitHub: Code
Huggingface: SciScore
Huggingface: Science-T2I Benchmark
Citation
@misc{li2025sciencet2iaddressingscientificillusions,
title={Science-T2I: Addressing Scientific Illusions in Image Synthesis},
author={Jialuo Li and Wenhao Chai and Xingyu Fu and Haiyang Xu and Saining Xie},
year={2025},
eprint={2504.13129},
archivePrefix={arXiv},
primaryClass={cs.CV}… See the full description on the dataset page: https://huggingface.co/datasets/Jialuo21/Science-T2I.Science-T2I-C
Science-T2I-C Benchmark
Resources
Website
arXiv: Paper
GitHub: Code
Huggingface: SciScore
Huggingface: Science-T2I-Trainset
Benchmark Collection and Processing
Science-T2I-C is generated using the identical procedure as the training data, with a key adjustment to the prompts. This test set pushes the model further by introducing more intricate scenarios, incorporating contextual details like specific scene settings and diverse situations. Prompts in… See the full description on the dataset page: https://huggingface.co/datasets/Jialuo21/Science-T2I-C.geo170k-8k-r1-VisualPuzzles-TQA-ai2d-r1-RL-lmms-ScienceQA-IMG-A-OKVQAScienceQA-IMGScience-T2I-S
Science-T2I-S Benchmark
Resources
Website
arXiv: Paper
GitHub: Code
Huggingface: SciScore
Huggingface: Science-T2I-S&C Benchmark
Benchmark Collection and Processing
Science-T2I-S is generated using the identical procedure as the training data, ensuring a close match in stylistic and structural characteristics. This test set prioritizes simplicity by concentrating on well-defined regions, allowing for a focused evaluation of a model's performance on… See the full description on the dataset page: https://huggingface.co/datasets/Jialuo21/Science-T2I-S.ViRL39K-GradeSchool__ScienceScienceQA
Dataset Card Creation Guide
Dataset Summary
Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Supported Tasks and Leaderboards
Multi-modal Multiple Choice
Languages
English
Dataset Structure
Data Instances
Explore more samples here.
{'image': Image,
'question': 'Which of these states is farthest north?',
'choices': ['West Virginia', 'Louisiana', 'Arizona'… See the full description on the dataset page: https://huggingface.co/datasets/Daisyamanda/ScienceQA.scienceqa-problems-dataset-testScienceQA_RS_think
ScienceQA — ScienceQA_RS_think
Rejection-sampled from the ScienceQA train split. This split holds the accepted items, with the model's reasoning trace.
rows
4,630
QA pairs
13,834
shards
12
accepted / rejected (whole family)
13,834 / 641
accept rate
95.6%
verifier
anls
How the data was produced
A VLM answers every question at temperature 0 with reasoning enabled. Its answer is compared with
the official ground truth by the verifier… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/ScienceQA_RS_think.R1-Vision-ScienceQA
R1-Vision: Let's first take a look at the image
[🤗 Cold-Start Dataset] [📜 Report (Coming Soon)]
DeepSeek-R1 demonstrates outstanding reasoning abilities when tackling math, coding, puzzle, and science problems, as well as responding to general inquiries. However, as a text-only reasoning model, R1 cannot process multimodal inputs like images, which limits its practicality in certain situations. Exploring the potential for multimodal reasoning is an intriguing… See the full description on the dataset page: https://huggingface.co/datasets/yuyq96/R1-Vision-ScienceQA.ScienceQA_rejected
ScienceQA — ScienceQA_rejected
Rejection-sampled from the ScienceQA train split. This split holds the rejected items — the answer field holds the official ground truth.
rows
199
QA pairs
641
shards
2
accepted / rejected (whole family)
13,834 / 641
accept rate
95.6%
verifier
anls
The rejected split is training data, not just diagnostics: answer is the official ground truth, and wrong_vlm records what the model said instead.
How the data… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/ScienceQA_rejected.
