datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sciq
Dataset Card for "sciq"
Dataset Summary
The SciQ dataset contains 13,679 crowdsourced science exam questions about Physics, Chemistry and Biology, among others. The questions are in multiple-choice format with 4 answer options each. For the majority of the questions, an additional paragraph with supporting evidence for the correct answer is provided.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed… See the full description on the dataset page: https://huggingface.co/datasets/allenai/sciq.ScienceQA
Dataset Card Creation Guide
Dataset Summary
Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Supported Tasks and Leaderboards
Multi-modal Multiple Choice
Languages
English
Dataset Structure
Data Instances
Explore more samples here.
{'image': Image,
'question': 'Which of these states is farthest north?',
'choices': ['West Virginia', 'Louisiana', 'Arizona', 'Oklahoma'],
'answer': 0… See the full description on the dataset page: https://huggingface.co/datasets/derek-thomas/ScienceQA.SciCodeThis dataset was presented in SciCode: A Research Coding Benchmark Curated by Scientists.
SciTS
SciTS: Scientific Time Series Understanding and Generation with LLMs
This repository contains the official dataset for SciTS: Scientific Time Series Understanding and Generation with LLMs (ICLR 2026). SciTS is a large-scale benchmark designed to evaluate the capabilities of large language models on complex scientific time series data. It spans 12 scientific disciplines, 43 distinct tasks, and includes 54,023 instances.
Dataset Structure
The benchmark is organized… See the full description on the dataset page: https://huggingface.co/datasets/ZJTustc/SciTS.SciTS
SciTS: Scientific Time Series Understanding and Generation with LLMs
This repository contains the official dataset for SciTS: Scientific Time Series Understanding and Generation with LLMs (ICLR 2026). SciTS is a large-scale benchmark designed to evaluate the capabilities of large language models on complex scientific time series data. It spans 12 scientific disciplines, 43 distinct tasks, and includes 54,023 instances.
Dataset Structure
The benchmark is organized… See the full description on the dataset page: https://huggingface.co/datasets/OpenTSLab/SciTS.xlam-function-calling-60k-raw
XLAM Function Calling 60k Raw Dataset
This dataset includes train and test splits derived from Salesforce/xlam-function-calling-60k.
Train split size: 95% of the original dataset
Test split size: 5% of the original dataset
SciKnowEval
SciKnowEval
Evaluating Multi-level Scientific Knowledge of Large Language Models
Please refer to our repository and paper for more details.
博学之 ,审问之 ,慎思之 ,明辨之 ,笃行之。
—— 《礼记 · 中庸》 Doctrine of the Mean
The Scientific Knowledge Evaluation (SciKnowEval) benchmark for Large Language Models (LLMs) is inspired by the profound principles outlined in the “Doctrine of the Mean” from ancient Chinese philosophy. This benchmark is designed to assess LLMs based on their proficiency in… See the full description on the dataset page: https://huggingface.co/datasets/hicai-zju/SciKnowEval.SciMDR-Evalmath-code-science-deepseek-r1-en
R1 Dataset Collection
Aggregated high-quality English prompts and model-generated responses from DeepSeek R1 and DeepSeek R1-0528.
Dataset Summary
The R1 Dataset Collection combines multiple public DeepSeek-generated instruction-response corpora into a single, cleaned, English-only JSONL file. Each example consists of a <|user|> prompt and a <|assistant|> response in one "text" field. This release includes:
~21,000 examples from the DeepSeek-R1-0528 Distilled Custom… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/math-code-science-deepseek-r1-en.BenchMAX_Science
Dataset Sources
Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Link: https://huggingface.co/papers/2502.07346
Repository: https://github.com/CONE-MT/BenchMAX
Dataset Description
BenchMAX_Science is a dataset of BenchMAX, sourcing from GPQA, which evaluates the natural science reasoning capability in multilingual scenarios.
We extend the original English dataset to 16 non-English languages.
The data is first translated by Google… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Science.scicode_wiki
SciCode Wikipedia Concept Corpus
中文说明
这个数据集是围绕 SciCode benchmark 构造的 Wikipedia Markdown 知识语料。
数据分为五个相互独立的部分:
原始 seed 概念集:每道题的 core_concepts 和 adjacent_concepts。
首轮 expanded 概念集:使用 deepseek-v4-flash 为每道题扩展 25 个概念。
2026-06-12 expanded 概念集:综合 seed 概念和首轮扩展结果,使用
gpt-5.5 为每道题继续扩展 70 个不重复概念。
2026-07-02 expanded 概念集:综合此前所有概念,使用
deepseek-v4-pro 为每道题继续扩展 50 个更发散的概念,并完成 Wikipedia 抓取。
2026-07-02 partial expanded 概念集:使用 deepseek-v4-pro 继续扩展后,
上传截至 2026-07-03 16:16:26 +08:00… See the full description on the dataset page: https://huggingface.co/datasets/dabingzz/scicode_wiki.hle_material_science
HLE Material Science: A Specialized Benchmark for Materials Science
A Materials Science Subset of Humanity's Last Exam (HLE)
Overview
HLE Material Science is a carefully curated materials science subset derived from the Humanity's Last Exam (HLE) dataset, containing 106 high-quality expert-level questions covering 25+ materials science subfields, with 97% of questions rated as high confidence.
This dataset is designed to evaluate large language models'… See the full description on the dataset page: https://huggingface.co/datasets/TalentZHOU/hle_material_science.SciCode
Dataset Card for Dataset Name
Official Description (from the authors):
Since language models (LMs) now outperform average humans on many challenging tasks,
it has become increasingly difficult to develop challenging, high-quality, and realistic evaluations.
We address this issue by examining LMs' capabilities to generate code for solving real scientific research problems.
Incorporating input from scientists and AI researchers in 16 diverse natural science sub-fields,
including… See the full description on the dataset page: https://huggingface.co/datasets/akshathmangudi/SciCode.LimAgents_limitation_data_scientific_papers_with_cited_papers
LimAgents Data
This dataset contains scientific paper metadata and extracted limitation information prepared for use with LLM Agents.The data comes from NeurIPS 2021–2022 papers and related OpenReview reviews, enriched with Cited in and Cited by information.
Dataset Structure
The repository contains two main directories:
1. NeurIPS_21_22_Lim_OPR_with_cited_in_by_papers
This directory includes one JSON file per paper. Each file contains:
title: Original paper… See the full description on the dataset page: https://huggingface.co/datasets/iaadlab/LimAgents_limitation_data_scientific_papers_with_cited_papers.islamic-sciences
islamlab — The Islamic Sciences Corpus
The Islamic sciences other than Qur'an and hadith, as their authors wrote
them: 4,022 works by scholars who died between the
0st and the 14th Hijri century, cut along their own chapter
and biographical-entry boundaries into 1,864,389 units
(3.41 billion characters of Arabic), each carrying the volume and
page it sits on so a quotation can be cited rather than merely produced.
Scope is Ahl al-Sunnah wa'l-Jamāʿah, and the gate is the author… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/islamic-sciences.HiPhO
🥇 HiPhO: High School Physics Olympiad Benchmark
[🏆 Leaderboard]
[📊 Dataset]
[✨ GitHub]
[📄 Paper]
📊 New (Dec. 8): Results from Gemini-3-Pro, DeepSeek-V3.2-Speciale, and Kimi-K2-Thinking have been added to the HiPhO Leaderboard. Notably, Gemini-3-Pro achieved gold-medal performance across all 13 Olympiads in HiPhO.
🧩 New (Nov. 5):We added CPhO 2025 (Chinese Physics Olympiad) — the national final theoretical exam to the HiPhO benchmark.
🏆 New (Sep. 16): We launched "PhyArena", a… See the full description on the dataset page: https://huggingface.co/datasets/SciYu/HiPhO.SciEvo
🎓 SciEvo: A Longitudinal Scientometric Dataset
Best Paper Award at the 1st Workshop on Preparing Good Data for Generative AI: Challenges and Approaches (Good-Data @ AAAI 2025)
SciEvo is a large-scale dataset that spans over 30 years of academic literature from arXiv, designed to support scientometric research and the study of scientific knowledge evolution. By providing a comprehensive collection of over two million publications, including detailed metadata and citation… See the full description on the dataset page: https://huggingface.co/datasets/Ahren09/SciEvo.xlam-function-calling-60k-raw-augmented
XLAM Function Calling 60k Raw Augmented Dataset
This dataset includes augmented train and test splits derived from product-science/xlam-function-calling-60k-raw.
Train split size: Original size plus augmented data
Test split size: Original size plus augmented data
Augmentation Details
This dataset has been augmented by modifying function names in the original data. Randomly selected function names have underscores replaced with periods at random positions… See the full description on the dataset page: https://huggingface.co/datasets/product-science/xlam-function-calling-60k-raw-augmented.contextualizing-scientific-claimsThis repository hosts the training/dev datasets and evaluation scripts for the 2024 Workshop on Scholarly Document Processing Shared Task: Context24: Contextualizing Scientific Figures and Tables
Background and Problem
People read and use scientific claims both within the scientific process (e.g., in literature reviews, problem formulation, making sense of conflicting data) and outside of science (e.g., evidence-informed deliberation). When doing so, it is critical to contextualize… See the full description on the dataset page: https://huggingface.co/datasets/joelchan/contextualizing-scientific-claims.chinese-materials-science-open-intelligence
🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset
Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.scicode
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Zilinghan/scicode.SciFIBench
SciFIBench
Jonathan Roberts, Kai Han, Neil Houlsby, and Samuel Albanie
NeurIPS 2024
Note: This repo has been updated to add two splits ('General_Figure2Caption' and 'General_Caption2Figure') with an additional 1000 questions. The original version splits are preserved and have been renamed as follows: 'Figure2Caption' -> 'CS_Figure2Caption' and 'Caption2Figure' -> 'CS_Caption2Figure'.
Dataset Summary
The SciFIBench (Scientific Figure… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/SciFIBench.science_reasoning
science_reasoning
Mistral-7B의 과학 지식·추론 능력 향상을 위해 6개 공개 과학 객관식 QA 데이터셋을 통일 포맷으로 변환하고, ARC-Challenge test와의 오염을 제거한 데이터셋입니다.
원본 데이터셋
allenai/sciq
allenai/openbookqa (main)
allenai/qasc
allenai/quartz
allenai/ai2_arc (ARC-Easy / ARC-Challenge)
nguyen-brat/worldtree
전처리
포맷 통일: 각 데이터셋의 서로 다른 스키마를 unique_id, orig_id, source, question, choices, answer, support 필드로 변환. support는 근거 문단/문장으로, 데이터셋별 원본 필드(support/fact/para/cot)에서 구성하거나 없으면 빈 문자열.… See the full description on the dataset page: https://huggingface.co/datasets/seonjeongh/science_reasoning.SciCode-Programming-Problems
DATA3: Programming Problems Generation Dataset
Dataset Overview
DATA3 is a large-scale programming problems generation dataset that contains AI-generated programming problems inspired by real scientific computing code snippets. The dataset consists of 22,532 programming problems, each paired with a comprehensive solution. These problems focus on scientific computing concepts such as numerical algorithms, data analysis, mathematical modeling, and computational methods in… See the full description on the dataset page: https://huggingface.co/datasets/SciCodePile/SciCode-Programming-Problems.EntLORE
🏛️ EntLORE
A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering
🤗 Paper page · 📄 arXiv:2608.10679 · 💻 Code & baselines
TL;DR — Enterprise answers often depend on organizational relations that no single document
states. EntLORE reconstructs an audited enterprise truth graph, releases an anonymized document
world whose gold answers and proofs are computed from that graph, and withholds the target
derived… See the full description on the dataset page: https://huggingface.co/datasets/Akrin-scitix/EntLORE.Science-QnA
Science-QnA
The Science-QnA is a large-scale, high-quality science-focused dataset (~5.63M rows) curated using synthetic data generation through distillation techniques and select open-source resources. Designed to train and evaluate reasoning-capable models in science domains with emphasis on conceptual understanding, numerical problem-solving, and exam-style Q&A patterns across Physics, Chemistry, Biology, and Mathematics.
Summary
• Domain: Science, Physics… See the full description on the dataset page: https://huggingface.co/datasets/169Pi/Science-QnA.Dr.SCI
Dr. SCI Dataset (Reproduced)
[📜 Original Paper] •
[🤗 Reproduced Dataset] •
[💻 Reproduced Github]
Disclaimer: This is an unofficial reproduction of the Dr. SCI dataset introduced in"Improving Data and Reward Design for Scientific Reasoning in Large Language Models" [arXiv].A detailed implementation of the curation process is available in my GitHub Repo.This work is not affiliated with or endorsed by the original authors. Please refer to the original paper for… See the full description on the dataset page: https://huggingface.co/datasets/MiniByte-666/Dr.SCI.hle_material_science
HLE Material Science: A Specialized Benchmark for Materials Science
A Materials Science Subset of Humanity's Last Exam (HLE)
Overview
HLE Material Science is a carefully curated materials science subset derived from the Humanity's Last Exam (HLE) dataset, containing 106 high-quality expert-level questions covering 25+ materials science subfields, with 97% of questions rated as high confidence.
This dataset is designed to evaluate large language models'… See the full description on the dataset page: https://huggingface.co/datasets/stonelight/hle_material_science.SciDocBench
SciDocBench
Official data release for SciDocBench: A Workflow-Centered Benchmark and Data
Pipeline for Scientific Document Understanding.
Project and evaluation code: InternLM/SciDocBench.
Dataset Summary
SciDocBench contains 496 evaluation instances derived from
124 expert-authored scientific-document questions.
Each question is represented under four matched settings that cross English/Chinese
questions with all-images-first/interleaved document representations.… See the full description on the dataset page: https://huggingface.co/datasets/HenryExcellent/SciDocBench.SciEGQA-Train
SciEGQA Training Set
SciEGQA is a scientific document question answering and reasoning dataset with semantic evidence grounding, where supporting evidence is represented as semantically coherent document regions annotated with bounding boxes.
Project page: https://yuwenhan07.github.io/SciEGQA-project/
Paper: https://arxiv.org/abs/2511.15090
Repository: https://github.com/yuwenhan07/SciEGQA
Quick start
import json
from pathlib import Path
from PIL import Image… See the full description on the dataset page: https://huggingface.co/datasets/Yuwh07/SciEGQA-Train.
