datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ScienceQA
Dataset Card Creation Guide
Dataset Summary
Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Supported Tasks and Leaderboards
Multi-modal Multiple Choice
Languages
English
Dataset Structure
Data Instances
Explore more samples here.
{'image': Image,
'question': 'Which of these states is farthest north?',
'choices': ['West Virginia', 'Louisiana', 'Arizona', 'Oklahoma'],
'answer': 0… See the full description on the dataset page: https://huggingface.co/datasets/derek-thomas/ScienceQA.ScienceQA
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of derek-thomas/ScienceQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@inproceedings{lu2022learn,
title={Learn to Explain: Multimodal Reasoning via Thought… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/ScienceQA.SciMMIR
Dataset Card for "SciMMIR_dataset"
SciMMIR
This is the repo for the paper SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval.
In this paper, we propose a novel SciMMIR benchmark and a corresponding dataset designed to address the gap in evaluating multi-modal information retrieval (MMIR) models in the scientific domain.
It is worth mentioning that we define a data hierarchical architecture of "Two subsets, Five subcategories" and use human-created… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/SciMMIR.SciMDR-EvalS1-MMAlignS1-MMAlign
A Large-Scale Multi-Disciplinary Scientific Multimodal Dataset
S1-MMAlign is a large-scale, multi-disciplinary multimodal dataset comprising over 15.5 million high-quality image-text pairs derived from 2.5 million open-access scientific papers.
Multimodal learning has revolutionized general domain tasks, yet its application in scientific discovery is hindered by the profound semantic gap between complex scientific imagery and sparse textual descriptions. S1-MMAlign aims to… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-MMAlign.science-datalake
Science Data Lake
A unified, portable science data lake integrating 7 scholarly datasets (~525 GB Parquet) with cross-dataset DOI normalization, 13 scientific ontologies (1.3M terms), and a reproducible ETL pipeline.
Note: One additional source (Semantic Scholar S2AG) is supported by the pipeline but is not redistributed here due to its API terms of service. See Not Included in This Upload below.
What's Unique
This dataset enables queries… See the full description on the dataset page: https://huggingface.co/datasets/J0nasW/science-datalake.vidore_v3_computer_scienceViDoRe V3 : Computer Science
This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark.
About ViDoRe v3
ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science.sciencemysterybench-transcriptsai-scientist-blog-assetsSciDraw-6K
SciDraw-6K: A Multilingual Scientific Illustration Dataset Generated by Google Gemini
Dataset Summary
SciDraw-6K is a curated dataset of 6,291 scientific illustrations synthesized by Google Gemini image-generation models, each paired with prompts in 11 languages (English, Chinese, Japanese, Korean, German, French, Spanish, Brazilian Portuguese, Traditional Chinese, Italian, and Russian).
Images span 8 broad scientific categories: biomedical, chemistry, materials… See the full description on the dataset page: https://huggingface.co/datasets/SciDrawAI/SciDraw-6K.ScienceQA-IMG
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted and filtered version of derek-thomas/ScienceQA with only image instances. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@inproceedings{lu2022learn,
title={Learn to Explain:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/ScienceQA-IMG.SciGraphQA-295K-train
Dataset Card for Dataset Name
Here is a filled out dataset card for the SciGraphQA dataset:
## Dataset Description
Homepage: https://github.com/findalexli/SciGraphQA
Repository: https://huggingface.co/datasets/alexshengzhili/SciGraphQA-295K-train
Paper: https://arxiv.org/abs/2308.03349
Leaderboard: N/A
Point of Contact Alex Li alex.shengzhi@gmail.com:
### Dataset Summary
SciGraphQA is a large-scale synthetic multi-turn question-answering dataset for scientific graphs. It… See the full description on the dataset page: https://huggingface.co/datasets/alexshengzhili/SciGraphQA-295K-train.self-contradictory
Introduction
Official dataset of the ECCV24 paper, "Dissecting Dissonance: Benchmarking Large Multimodal Models Against Self-Contradictory Instructions".
Website: https://selfcontradiction.github.io
Github: https://github.com/shiyegao/Self-Contradictory-Instructions-SCI
Sample usage
In the paper, “SCI-Core (1%), SCI-Base (10%), and SCI-All (100%)” denote the small, medium, and full splits of the Hugging Face dataset, respectively.
Language-Language
from… See the full description on the dataset page: https://huggingface.co/datasets/sci-benchmark/self-contradictory.SciFormaData-700KSciFormaData-700K: Training Data for Scientific Diagram Generation
SciFormaData-700K is the official training dataset for
SciForma. It contains scientific
methodology-diagram records collected from arXiv papers spanning January
2015–December 2025, structured generation prompts, multi-resolution training
targets, and axis-specific editing triplets.
Features
🧩 Structure-aware prompts. Detailed descriptions organize diagram
components… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/SciFormaData-700K.scidocbench-sft
scidocbench-sft
This repository contains a packaged export of the sft split of SciDocBench.
Original file: sft_0420_26k_max20_images.json
Records: 24674
Referenced images: 60246
All image paths in the dataset file have been rewritten from cluster-local absolute
paths to repository-relative paths under images/, so the dataset can be moved or
downloaded without depending on the original filesystem layout.
scin
SCIN Dataset
The SCIN (Skin Condition Image Network) open access dataset aims to supplement publicly available dermatology datasets from health system sources with representative images from internet users. To this end, the SCIN dataset was collected from Google Search users in the United States through a voluntary, consented image donation application. The SCIN dataset is intended for health education and research, and to increase the diversity of dermatology images available for… See the full description on the dataset page: https://huggingface.co/datasets/google/scin.SciFigPlag-Bench
SciFigPlag-Bench: A Benchmark for Provenance-Aware Scientific Figure Plagiarism Detection
🌐 Homepage |
📖 arXiv
SciFigPlag-Bench is a benchmark for provenance-aware scientific figure plagiarism detection. It evaluates whether a suspicious figure reuses evidence from a specific source figure, which figure is the original source, how the reused content has been transformed, and where the reused evidence appears.
The benchmark is designed to evaluate vision-language models… See the full description on the dataset page: https://huggingface.co/datasets/FreeLand123/SciFigPlag-Bench.vidore_v3_computer_science_mteb_format
Vidore3ComputerScienceRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Retrieve associated pages according to questions.
Task category
t2i
Domains
Academic
Reference
https://huggingface.co/blog/QuentinJG/introducing-vidore-v3
Source datasets:
vidore/vidore_v3_computer_science
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science_mteb_format.S1-Omni-Corpus-10K
S1-Omni-Corpus-10K
An open-source scientific multimodal reasoning dataset subset for S1-Omni
🧬 Model Introduction
S1-Omni is a unified scientific multimodal reasoning model for scientific understanding, prediction, and generation. It is developed by the ScienceOne AI team of the Chinese Academy of Sciences.
S1-Omni addresses fragmented scientific AI capabilities with a shared backbone for cross-disciplinary, cross-modal, and cross-task understanding and reasoning… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-Omni-Corpus-10K.TreeOil_Painting_ScientificJourney_Thailand_CaseStudy🧪 Tree Oil Painting: A Scientific Journey – Thailand Case Study
This dataset documents a rare and detailed forensic investigation of a mysterious 19th-century oil painting, known as The Tree Oil Painting, using scientific methods and AI-assisted analysis. Compiled in Thailand between 2015 and 2025, this work represents a grassroots effort to validate the painting’s origins through physical evidence, pigment mapping, synchrotron spectroscopy, and historical comparison.
🧩 Overview
Title: Tree… See the full description on the dataset page: https://huggingface.co/datasets/HaruthaiAi/TreeOil_Painting_ScientificJourney_Thailand_CaseStudy.scientific-figures-captions-context
Dataset Card for Scientific Figures, Captions, and Context
A novel vision-language dataset of scientific figures taken directly from research papers.
We scraped approximately ~150k papers, with about ~690k figures total. We extracted each figure's caption and label from the paper. In addition, we searched through each paper to find references of each figure and included the surrounding text as 'context' for this figure.
All figures were taken from arXiv research papers.… See the full description on the dataset page: https://huggingface.co/datasets/mawadalla/scientific-figures-captions-context.scientific-chart-qa-17k
Scientific Chart QA, 17,070 rows
A multimodal chart-interpretation dataset built around one idea: teaching a model when not to
answer matters as much as teaching it to answer.
One in seven questions here cannot be answered from its figure, and the correct response is
cannot be determined. Baseline vision-language models overwhelmingly guess a plausible-looking
number instead. That is the behaviour this set targets.
The four things worth… See the full description on the dataset page: https://huggingface.co/datasets/manifesta/scientific-chart-qa-17k.C4-Eval
C4-Eval
C4-Eval is the evaluation set for C4 Bench, a Chengyu-based benchmark for measuring whether multimodal language models can understand cross-concept creativity. The release contains the original images, the corresponding idiom answers, and the complete task-specific questions used for evaluation.
221 base items: 37 human-designed seed figures and 184 bridge-controlled synthetic figures.
1,105 evaluation instances: five task formulations for every base item.
Language:… See the full description on the dataset page: https://huggingface.co/datasets/sci-m-wang/C4-Eval.SciFigAlign
SciFigAlign
Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence
Paper · Code · Dataset
Scientific figure assessment in peer review is not natural-image IQA. A figure must be legible, support the manuscript’s claims, and present a clear visual hierarchy.
This dataset is the full SciFigAlign corpus: 3,857 figure crops from 3,126 ICLR / NeurIPS / ICML papers, labeled on four peer-review dimensions (1–5): Clarity, Relevance, Informativeness… See the full description on the dataset page: https://huggingface.co/datasets/haihanlamu/SciFigAlign.SciGenEdit-10K
SciGenEdit-10K
An Open Dataset for Scientific Image Generation and Editing
English | 简体中文
📖 Introduction
SciGenEdit-10K is a public subset released with the S1-Omni-Image project. It is designed for research on scientific image generation, scientific image editing, and multi-turn scientific image generation and editing.
S1-Omni-Image is a unified multimodal model developed by the ScienceOne team at the Chinese Academy of Sciences for scientific… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/SciGenEdit-10K.SciVisCap
Dataset Card for SciVisCap
Dataset Summary
SciVisCap is a dataset for captioning scientific visualization (SciVis)
figures -- generating a caption for a SciVis rendering, optionally using the
paragraph(s) in the source paper that reference it. It contains 3,539
SciVis figures from 992 IEEE Vis and SciVis papers, drawn from the VIS30K
corpus and paired with their published captions and figure-referencing
paragraphs.
Supported Tasks
Image captioning… See the full description on the dataset page: https://huggingface.co/datasets/PVIS2027-JT-6597/SciVisCap.ReasonEMThis dataset is prepared for NIPS2026 ED track.
colour-checker-detection-dataset
Colour - Checker Detection - Dataset
An image dataset of colour rendition charts.
This dataset is structured according to Ultralytics YOLO format and ready to use with YOLOv8.
The colour-science/colour-checker-detection-models models resulting from the YOLOv8 segmentation training are supporting colour rendition charts detection in the Colour Checker Detection Python package.
Classes
ColorCheckerClassic24: Calibrite / X-Rite ColorCheckerClassic 24
Contact &… See the full description on the dataset page: https://huggingface.co/datasets/colour-science/colour-checker-detection-dataset.SciVer
SCIVER: A Benchmark for Multimodal Scientific Claim Verification
🌐 Github •
📖 Paper •
🤗 Data
📰 News
[May 15, 2025] SciVer has been accepted by ACL 2025 Main!
👋 Overview
SCIVER is the first benchmark specifically designed to evaluate the ability of foundation models to verify scientific claims across text, charts, and tables. It challenges models to reason over complex, multimodal contexts with fine-grained entailment labels and… See the full description on the dataset page: https://huggingface.co/datasets/chengyewang/SciVer.Handwritten-Computer-Science-Notes-Dataset
English Handwritten Computer Science Notes Dataset
This dataset contains high-resolution images of handwritten computer science notes written in English. It includes algorithm explanations, code snippets, flowcharts, theoretical content, and annotations. The dataset is designed to support AI research in handwriting recognition, OCR, and document understanding specifically for computer science education.
Contact
For queries or collaborations related to this dataset… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Handwritten-Computer-Science-Notes-Dataset.
