datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ScienceQA
Dataset Card Creation Guide
Dataset Summary
Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Supported Tasks and Leaderboards
Multi-modal Multiple Choice
Languages
English
Dataset Structure
Data Instances
Explore more samples here.
{'image': Image,
'question': 'Which of these states is farthest north?',
'choices': ['West Virginia', 'Louisiana', 'Arizona', 'Oklahoma'],
'answer': 0… See the full description on the dataset page: https://huggingface.co/datasets/derek-thomas/ScienceQA.ScienceQA
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of derek-thomas/ScienceQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@inproceedings{lu2022learn,
title={Learn to Explain: Multimodal Reasoning via Thought… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/ScienceQA.xlam-function-calling-60k-raw
XLAM Function Calling 60k Raw Dataset
This dataset includes train and test splits derived from Salesforce/xlam-function-calling-60k.
Train split size: 95% of the original dataset
Test split size: 5% of the original dataset
ScienceQA_text_only
Dataset Card for "scienceQA_text_only"
ScienceQA text-only examples (examples where no image was initially present, which means they should be doable with text-only models.)
@article{10.1007/s00799-022-00329-y,
author = {Saikh, Tanik and Ghosal, Tirthankar and Mittal, Amish and Ekbal, Asif and Bhattacharyya, Pushpak},
title = {ScienceQA: A Novel Resource for Question Answering on Scholarly Articles},
year = {2022},
journal = {Int. J. Digit. Libr.},
month = {sep}
}
mmu_manga
mmu_manga HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_manga.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be installed via… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_manga.vidore_v3_computer_scienceViDoRe V3 : Computer Science
This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark.
About ViDoRe v3
ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science.judged_science_completionsseverity_ablation_scienceScienceQA-IMG
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted and filtered version of derek-thomas/ScienceQA with only image instances. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@inproceedings{lu2022learn,
title={Learn to Explain:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/ScienceQA-IMG.French-Science-Commons
French Science Commons
French Science Commons (Commun numérique des sciences en français) rassemble des publications scientifiques d'origine française en accès ouvert, couvrant une période de vingt ans, de 2007 à 2026. Il comprend 1 248 860 documents scientifiques — 1 189 628 articles et 59 232 thèses — indexés à travers de multiples dépôts académiques en accès public, tels que HAL, OpenAlex, des revues scientifiques, des dépôts institutionnels, et d'autres.
Le corpus est conçu… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-Science-Commons.vidore_v3_computer_science_mteb_format
Vidore3ComputerScienceRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Retrieve associated pages according to questions.
Task category
t2i
Domains
Academic
Reference
https://huggingface.co/blog/QuentinJG/introducing-vidore-v3
Source datasets:
vidore/vidore_v3_computer_science
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science_mteb_format.pseudolabel-science-large-v3-timestamp
Pseudolabel science context audio using Whisper Large V3
Original audio from malaysia-ai/science-context-youtube, we split every 30 seconds and pseudolabelled using Whisper Large V3.
how to prepare the dataset
huggingface-cli download --repo-type dataset \
--include 'science-chunk-*.zip' \
--local-dir './' \
--max-workers 20 \
mesolitica/pseudolabel-science-large-v3-timestamp
wget… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/pseudolabel-science-large-v3-timestamp.hle_material_science
HLE Material Science: A Specialized Benchmark for Materials Science
A Materials Science Subset of Humanity's Last Exam (HLE)
Overview
HLE Material Science is a carefully curated materials science subset derived from the Humanity's Last Exam (HLE) dataset, containing 106 high-quality expert-level questions covering 25+ materials science subfields, with 97% of questions rated as high confidence.
This dataset is designed to evaluate large language models'… See the full description on the dataset page: https://huggingface.co/datasets/TalentZHOU/hle_material_science.science_studies_textbooknatural-science-reasoning
Natural Sciences Reasoning: the "smolest" reasoning dataset
A smol-scale open dataset for reasoning tasks using Hugging Face Inference Endpoints. While intentionally limited in scale, this resource prioritizes:
Reproducible pipeline for reasoning tasks using a variety of models (Deepseek V3, Deepsek-R1, Llama70B-Instruct, etc.)
Knowledge sharing for domains other than Math and Code reasoning
In this repo, you can find:
The prompts and the pipeline (see the config file).
The… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/natural-science-reasoning.openthoughts_science_1kislamic-sciences
islamlab — The Islamic Sciences Corpus
The Islamic sciences other than Qur'an and hadith, as their authors wrote
them: 4,022 works by scholars who died between the
0st and the 14th Hijri century, cut along their own chapter
and biographical-entry boundaries into 1,864,389 units
(3.41 billion characters of Arabic), each carrying the volume and
page it sits on so a quotation can be cited rather than merely produced.
Scope is Ahl al-Sunnah wa'l-Jamāʿah, and the gate is the author… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/islamic-sciences.open-thoughts-sciencefineweb-1m-sampleolmocr_science_pdfs-literaturehttps://huggingface.co/datasets/allenai/dolma3_pool/tree/main/data/olmocr_science_pdfs-literature
EU-Science-Commonsqwq_mix_qwen3_scienceQwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179
mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
60.7
90.8
89.4
63.2
52.4
48.5
27.4
26.2
48.3
12.0
34.3
34.7
AIME24
Average Accuracy: 60.67% ± 2.25%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179.RaR-Science
Dataset Summary
RaR-Science is a dataset curated for training and evaluating language models on science domain using structured rubric-based supervision. Each example includes a science related question, a reference answer, and checklist-style rubric annotations generated via OpenAI's o3-mini. This dataset is introduced in Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains.
Dataset Structure
Data Fields
Each example contains:… See the full description on the dataset page: https://huggingface.co/datasets/anisha2102/RaR-Science.xlam-function-calling-60k-raw-augmented
XLAM Function Calling 60k Raw Augmented Dataset
This dataset includes augmented train and test splits derived from product-science/xlam-function-calling-60k-raw.
Train split size: Original size plus augmented data
Test split size: Original size plus augmented data
Augmentation Details
This dataset has been augmented by modifying function names in the original data. Randomly selected function names have underscores replaced with periods at random positions… See the full description on the dataset page: https://huggingface.co/datasets/product-science/xlam-function-calling-60k-raw-augmented.science_traces_original_DeepSeek-R1-Distill-Qwen-32BQwen2.5-7B-Instruct_qwq_mix_r1_science_eval_2870
mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_2870
Precomputed model outputs for evaluation.
Evaluation Results
AIME24
Average Accuracy: 60.67% ± 2.20%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
70.00%
21
30
2
53.33%
16
30
3
53.33%
16
30
4
66.67%
20
30
5
63.33%
19
30
6
66.67%
20
30
7
60.00%
18
30
8
46.67%
14
30
9
63.33%
19
30
10
63.33%
19
30
mmu_apogee_dr17
mmu_apogee_dr17 HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_apogee_dr17.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_apogee_dr17.llama-nemotron-science-reasoning-on-canonical-think-full
Llama-Nemotron science reasoning — Delphi canonical-think (COMPLETE, no length filter)
The complete reasoning:on science split of
nvidia/Llama-Nemotron-Post-Training-Dataset, converted once into the canonical
Delphi chat-template thinking format. 708,920 rows.
Unlike the cold-start warmup slice
open-athena/llama-nemotron-science-reasoning-on-le3000tok-100k
(and its -canonical-think variant), this build applies no length cap and no subsample — every
long-CoT science example is… See the full description on the dataset page: https://huggingface.co/datasets/laion/llama-nemotron-science-reasoning-on-canonical-think-full.mmu_hsc_pdr3_wide_21
mmu_hsc_pdr3_wide_21 HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_hsc_pdr3_wide_21.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_hsc_pdr3_wide_21.
