science
Datasets
All datasets matching “science”terminal-bench-science
Terminal-Bench-Science
The primary source is hosted on GitHub, please open
issues and pull requests there, not here.
Terminal-Bench-Science is a benchmark of real-world computational research
workflows across the life, physical, earth, mathematical, and engineering sciences. Like Terminal-Bench, it's a
continuous benchmark: releases are published as tags on the source repo. This dataset mirrors that on the Hub:
one repo, tagged per version, instead of a separate repo per… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-science.ScienceQA
Dataset Card Creation Guide
Dataset Summary
Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Supported Tasks and Leaderboards
Multi-modal Multiple Choice
Languages
English
Dataset Structure
Data Instances
Explore more samples here.
{'image': Image,
'question': 'Which of these states is farthest north?',
'choices': ['West Virginia', 'Louisiana', 'Arizona', 'Oklahoma'],
'answer': 0… See the full description on the dataset page: https://huggingface.co/datasets/derek-thomas/ScienceQA.terminal-bench-science-lfs
Terminal-Bench-Science — task input mirror
Large input files for Terminal-Bench-Science
tasks, which cannot be committed to git. Tasks pull from here at container build
time, pinned to a commit SHA and verified against a checksum file that ships in
the task directory.
One top-level prefix per task; everything lives under <task-name>/input/.
Benchmark contamination canary
This dataset is benchmark material. If you are assembling a training corpus,
exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-science-lfs.ScienceQA
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of derek-thomas/ScienceQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@inproceedings{lu2022learn,
title={Learn to Explain: Multimodal Reasoning via Thought… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/ScienceQA.mmu_legacysurvey_dr10_south_21
mmu_legacysurvey_dr10_south_21 HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_legacysurvey_dr10_south_21.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_legacysurvey_dr10_south_21.Nemotron-SFT-Science-v2
Dataset Description:
Nemotron-Science-v2 is a science reasoning dataset with synthetic (synthetic MCQ, RQA) and non-synthetic vendor problems and LLM-generated solutions. It comprises three domains (Physics, Biology, and Chemistry), two question formats (multiple-choice questions [MCQ] and open questions [OpenQ]), and three generation setups: chain-of-thought (CoT) reasoning without tools, Python tool usage, and search tools usage with the Tavily API.
The solutions were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Science-v2.
