datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
terminal-bench-science
Terminal-Bench-Science
The primary source is hosted on GitHub, please open
issues and pull requests there, not here.
Terminal-Bench-Science is a benchmark of real-world computational research
workflows across the life, physical, earth, mathematical, and engineering sciences. Like Terminal-Bench, it's a
continuous benchmark: releases are published as tags on the source repo. This dataset mirrors that on the Hub:
one repo, tagged per version, instead of a separate repo per… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-science.ScienceQA
Dataset Card Creation Guide
Dataset Summary
Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Supported Tasks and Leaderboards
Multi-modal Multiple Choice
Languages
English
Dataset Structure
Data Instances
Explore more samples here.
{'image': Image,
'question': 'Which of these states is farthest north?',
'choices': ['West Virginia', 'Louisiana', 'Arizona', 'Oklahoma'],
'answer': 0… See the full description on the dataset page: https://huggingface.co/datasets/derek-thomas/ScienceQA.terminal-bench-science-lfs
Terminal-Bench-Science — task input mirror
Large input files for Terminal-Bench-Science
tasks, which cannot be committed to git. Tasks pull from here at container build
time, pinned to a commit SHA and verified against a checksum file that ships in
the task directory.
One top-level prefix per task; everything lives under <task-name>/input/.
Benchmark contamination canary
This dataset is benchmark material. If you are assembling a training corpus,
exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-science-lfs.ScienceQA
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of derek-thomas/ScienceQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@inproceedings{lu2022learn,
title={Learn to Explain: Multimodal Reasoning via Thought… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/ScienceQA.mmu_legacysurvey_dr10_south_21
mmu_legacysurvey_dr10_south_21 HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_legacysurvey_dr10_south_21.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_legacysurvey_dr10_south_21.Nemotron-SFT-Science-v2
Dataset Description:
Nemotron-Science-v2 is a science reasoning dataset with synthetic (synthetic MCQ, RQA) and non-synthetic vendor problems and LLM-generated solutions. It comprises three domains (Physics, Biology, and Chemistry), two question formats (multiple-choice questions [MCQ] and open questions [OpenQ]), and three generation setups: chain-of-thought (CoT) reasoning without tools, Python tool usage, and search tools usage with the Tavily API.
The solutions were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Science-v2.xlam-function-calling-60k-raw
XLAM Function Calling 60k Raw Dataset
This dataset includes train and test splits derived from Salesforce/xlam-function-calling-60k.
Train split size: 95% of the original dataset
Test split size: 5% of the original dataset
science-theory-textbooksarc-aphasia-bids
Aphasia Recovery Cohort (ARC)
Multimodal neuroimaging dataset for stroke-induced aphasia research.
Dataset Summary
The Aphasia Recovery Cohort (ARC) is a large-scale, longitudinal neuroimaging dataset containing multimodal MRI scans from 230 chronic stroke patients with aphasia. This HuggingFace-hosted version provides direct Python access to the BIDS-formatted data with embedded NIfTI files.
Metric
Count
Subjects
230
Sessions
902
T1-weighted scans
444… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/arc-aphasia-bids.ScienceQA_text_only
Dataset Card for "scienceQA_text_only"
ScienceQA text-only examples (examples where no image was initially present, which means they should be doable with text-only models.)
@article{10.1007/s00799-022-00329-y,
author = {Saikh, Tanik and Ghosal, Tirthankar and Mittal, Amish and Ekbal, Asif and Bhattacharyya, Pushpak},
title = {ScienceQA: A Novel Resource for Question Answering on Scholarly Articles},
year = {2022},
journal = {Int. J. Digit. Libr.},
month = {sep}
}
legacy-satvision-sr-pretrain-small
Satvision Pretraining Dataset - Small
Developed by: NASA GSFC CISTO Data Science Group
Model type: Pre-trained visual transformer model
License: Apache license 2.0
This dataset repository houses the pretraining data for the Satvision pretrained transformers.
This dataset was constructed using webdatasets to
limit the number of inodes used in HPC systems with limited shared storage. Each file has 100000
tiles, with pairs of image input and annotation. The data has been further… See the full description on the dataset page: https://huggingface.co/datasets/nasa-cisto-data-science-group/legacy-satvision-sr-pretrain-small.isles24-stroke
ISLES'24 Stroke Training Dataset
Multi-center longitudinal multimodal acute ischemic stroke training dataset from the ISLES'24 Challenge.
Overview
149 acute ischemic stroke training cases with:
Admission imaging (ses-01): Non-contrast CT, CT angiography, 4D CT perfusion
Follow-up imaging (ses-02): Post-treatment MRI (DWI, ADC)
Clinical data: Demographics, patient history, admission NIHSS, 3-month mRS outcomes
Annotations: Infarct masks, large vessel occlusion masks… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/isles24-stroke.SWE-bench-Science
SWE-bench Science
SWE-bench Science evaluates coding agents on software-engineering tasks drawn from scientific-computing repositories. The release contains 119 tasks across 20 scientific domains, with isolated environments and separate programmatic verifiers.
GitHub release repository: OpenMOSS/SWE-bench-Science
Runtime images: Docker Hub, pinned by immutable linux/amd64 digests
Evaluation framework: Pier, compatible with Harbor task format
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/SWE-bench-Science.S1-MMAlignS1-MMAlign
A Large-Scale Multi-Disciplinary Scientific Multimodal Dataset
S1-MMAlign is a large-scale, multi-disciplinary multimodal dataset comprising over 15.5 million high-quality image-text pairs derived from 2.5 million open-access scientific papers.
Multimodal learning has revolutionized general domain tasks, yet its application in scientific discovery is hindered by the profound semantic gap between complex scientific imagery and sparse textual descriptions. S1-MMAlign aims to… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-MMAlign.G1-Pretrain-Data-RepoScienceAgentBench
ScienceAgentBench
Update 04/30/2026: To mitigate false negatives in evaluation, we have released a verified version of ScienceAgentBench. Please load our benchmark using the following code going forward and make sure you follow the latest instructions in our github repository:
from datasets import load_dataset
ds = load_dataset("osunlp/ScienceAgentBench", split="verified")
The advancements of language language models (LLMs) have piqued growing interest in developing… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/ScienceAgentBench.Nemotron-Science-v1
Dataset Description:
Nemotron-Science-v1 is a synthetic science reasoning dataset with two subsets: an MCQA set that improves on the STEM portion of Nemotron-Post-Training-v1 using GPT-OSS-120B to generate GPQA-style questions and reasoning traces, and an RQA set of synthetic chemistry questions.
This dataset is ready for commercial use.
The Nemotron-Science-v1 dataset contains the following subsets:
MCQA
This subset is an improvement of the STEM subset in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Science-v1.mmu_manga
mmu_manga HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_manga.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be installed via… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_manga.science-datalake
Science Data Lake
A unified, portable science data lake integrating 7 scholarly datasets (~525 GB Parquet) with cross-dataset DOI normalization, 13 scientific ontologies (1.3M terms), and a reproducible ETL pipeline.
Note: One additional source (Semantic Scholar S2AG) is supported by the pipeline but is not redistributed here due to its API terms of service. See Not Included in This Upload below.
What's Unique
This dataset enables queries… See the full description on the dataset page: https://huggingface.co/datasets/J0nasW/science-datalake.Aneumo
Aneumo Datasets
AneumoDataset is a comprehensive multi-modal dataset containing 3D anatomical structures and simulated hemodynamic parameters for cerebral aneurysms, supporting both computational modeling and AI-based analysis.
science-activationsvidore_v3_computer_scienceViDoRe V3 : Computer Science
This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark.
About ViDoRe v3
ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science.Computer-Science
מאגר הנתונים CS26 HIT — מדעי המחשב
מאגר זה משמש לאחסון מרכזי של נתוני לימוד ומשאבים אקדמיים עבור סטודנטים למדעי המחשב במכון הטכנולוגי חולון. המידע המצוי כאן מונגש בצורה נוחה באמצעות פורטל גישה נפרד המאפשר ניווט ויזואלי וחיפוש יעיל בתוך התיקיות השונות.
קישורים וגישה
ניתן להשתמש בפורטל בכתובת https://cs26-cs26-portal.hf.space/
תנאי שימוש וזכויות יוצרים
כל חומרי הלימוד והתכנים המופיעים במאגר זה פתוחים וחופשיים לשימוש לצורכי למידה בלבד. ניתן לקחת את… See the full description on the dataset page: https://huggingface.co/datasets/CS26/Computer-Science.sciencemysterybench-transcriptsscience_textbooksrocket-science
Rocket Science
Time-aligned video, keyboard actions, game events, and per-frame game state for all four players, captured from a 2v2 Rocket League match.
This is the dataset behind MIRA, a real-time multiplayer world model trained to simulate Rocket League gameplay — by General Intuition and Kyutai, in collaboration with Epic Games. Code, technical report,
and a live demo: github.com/mira-wm/mira ·
mira-wm.com.
Loading
from mira.data import RocketScienceDataset… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/rocket-science.AI4FA-Diabimmune
Food Allergy Microbiome Dataset (Experimental)
Dataset Summary
This is an experimental microbiome dataset designed for exploratory research in food allergy classification. The dataset contains multiple data modalities (DNA embeddings, microbiome embeddings, raw DNA sequences) collected longitudinally at several timepoints.
Warning: This dataset is experimental. Its structure is frozen for ongoing research, and it is not ready for benchmarking.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/AI4FA-Diabimmune.judged_science_completionsseverity_ablation_scienceScienceQA-IMG
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted and filtered version of derek-thomas/ScienceQA with only image instances. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@inproceedings{lu2022learn,
title={Learn to Explain:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/ScienceQA-IMG.
