datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
islamic-sciences
islamlab — The Islamic Sciences Corpus
The Islamic sciences other than Qur'an and hadith, as their authors wrote
them: 4,022 works by scholars who died between the
0st and the 14th Hijri century, cut along their own chapter
and biographical-entry boundaries into 1,864,389 units
(3.41 billion characters of Arabic), each carrying the volume and
page it sits on so a quotation can be cited rather than merely produced.
Scope is Ahl al-Sunnah wa'l-Jamāʿah, and the gate is the author… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/islamic-sciences.SciEvo
🎓 SciEvo: A Longitudinal Scientometric Dataset
Best Paper Award at the 1st Workshop on Preparing Good Data for Generative AI: Challenges and Approaches (Good-Data @ AAAI 2025)
SciEvo is a large-scale dataset that spans over 30 years of academic literature from arXiv, designed to support scientometric research and the study of scientific knowledge evolution. By providing a comprehensive collection of over two million publications, including detailed metadata and citation… See the full description on the dataset page: https://huggingface.co/datasets/Ahren09/SciEvo.chinese-materials-science-open-intelligence
🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset
Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.scientific-verification
Scientific Verification Benchmark: NMC Cathodes
Dataset summary
The benchmark contains 50 scientific claims about NMC (lithium nickel manganese cobalt oxide) battery cathodes. Each claim is answered by Claude Opus 5, GPT 5.6 Luna and Gemini 3.1 Pro using a set of 20 open-access papers, producing 150 scored answers. The accompanying reference set contains 1,991 experiment-grounded measurements curated from 227 open-access papers, with experimental conditions and… See the full description on the dataset page: https://huggingface.co/datasets/GenData-Research/scientific-verification.SciDQA
SciDQA: A Deep Reading Comprehension Dataset over Scientific Papers
📄 Paper | 💻 Code
Scientific literature is typically dense, requiring significant background knowledge and deep comprehension for effective engagement. We introduce SciDQA, a new dataset for reading comprehension that challenges LLMs for a deep understanding of scientific articles, consisting of 2,937 QA pairs. Unlike other scientific QA datasets, SciDQA sources questions from peer reviews by domain experts and… See the full description on the dataset page: https://huggingface.co/datasets/yale-nlp/SciDQA.nasa-science-repos-sme-benchmark
NASA Science Repos SME Benchmark
A benchmark dataset for evaluating retrieval systems on NASA science repository discovery tasks. This dataset contains expert queries, a corpus of NASA science GitHub repositories, and relevance judgments.
Dataset Structure
Files
├── corpus.jsonl # 5,264 repositories with full metadata
├── queries.jsonl # 219 expert queries
└── qrels/
├── earth.tsv # Earth Science relevance judgments (162)
├──… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-repos-sme-benchmark.Global-Ocean-Science-Corpus
🌊 Global-Ocean-Science-Corpus (v2.0 Curated & Cleaned)
A Highly Curated, Large-Scale Pre-Training & RAG Corpus for Deep Ocean Sciences, Marine Biology, and Oceanography
Language Note: This dataset is a 100% English-language scientific corpus (language: "en") aggregating peer-reviewed literature, deep-sea exploration dossiers, and technical oceanographic reports from leading global marine institutes.
Global-Ocean-Science-Corpus, derin okyanus bilimleri, deniz biyolojisi… See the full description on the dataset page: https://huggingface.co/datasets/tilikumotp/Global-Ocean-Science-Corpus.sci-agent-verification-cascade
Scientific Agent Verification Cascade
Public evaluation fixtures and verified aggregate results for testing whether
scientific claims keep their source, meaning, uncertainty, and verification
requirements as they move between AI agents.
This dataset accompanies the
Scientific Agent Verification Cascade
codebase. Version 0.2.0
contains synthetic evaluation data and aggregate-only results. It contains no
raw hosted-model response, private holdout identifier,
source-record… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/sci-agent-verification-cascade.SciCloze-900
SciCloze-900
SciCloze-900 is a cloze-style GCSE Combined Science benchmark for evaluating small base language models.
The benchmark contains 900 multiple-choice cloze items:
300 biology items
300 chemistry items
300 physics items
Each item is designed as a natural text continuation rather than an instruction-style question. The intended evaluation method is to score each answer choice by average log probability per token as a continuation of the prompt.
Splits… See the full description on the dataset page: https://huggingface.co/datasets/veyra-ai/SciCloze-900.synthetic-science-v2-sample
Synthetic Scientific Research Threads — v2 (sample)
A synthetic continual-learning benchmark: each episode is a coherent sequence of
short fictional scientific research documents about a single made-up entity, with
per-document QA anchors. Later documents build on, revise, or supersede earlier
ones. Designed to stress test-time / meta-learning approaches where a model must
adapt to a stream of documents and answer questions grounded in what it has just
seen.
This is a sample… See the full description on the dataset page: https://huggingface.co/datasets/HerrHruby/synthetic-science-v2-sample.IEEE2026_BigData_MAS-4-Science-Matching
SciAgentTrace
An execution-layer trace resource for scientific-agent workload characterization.
A protocol fixes who reasons, what each role can see, when feedback returns, and
when a workflow stops. Those choices determine the sequence of model requests
that produces an answer, so protocol design is also workload design. Two
workflows that consume similar token totals can issue very different request
sequences. SciAgentTrace records that difference.
The matched core runs the… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/IEEE2026_BigData_MAS-4-Science-Matching.NCERT_Science_10threwrite-questions-real-words-sciency
real_words_sciency.csv - Question Rewriting Dataset
This dataset contains question rewriting outputs from the file real_words_sciency.csv.
Dataset Structure
The dataset contains the following columns:
custom_id: Unique identifier for each question
style: Rewriting style applied (e.g., "gibberish")
index: Numerical index
original: Original question text
rewritten: Rewritten version of the question
options: Multiple choice options (list format)
correct: Index of the… See the full description on the dataset page: https://huggingface.co/datasets/NLie2/rewrite-questions-real-words-sciency.science_leaderboard_submissionThis dataset contains the results used for Science Leaderboard
NCERT_Science_8thNCERT_Political_Science_12thsynthlabs-GLM-5.2-Science
GLM-5.2 Science Synth Reasoning
Synthetic reasoning traces for science questions from the GLM-5.2 science dataset. Each record contains a complex scientific question with SYNTH-style reasoning and a generated answer.
Dataset Summary
33,014 records (605 dupes + 726 incomplete/truncated removed from 34,345 source)
33,014 reasoning turns (99.9% format compliance)
Average 3,094 chars per reasoning trace
Models Used
Model
Records… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/synthlabs-GLM-5.2-Science.NCERT_Science_6thPubmedFact1klicense: apache-2.0
pubid: Directly inherited from the input.
claim: Reformatted from the original "question" field to express a clear claim statement.
context: Contains:
contexts: A subset of the original context information.
labels: Corresponding labels related to the context.
final_decision: Converted from a textual decision to a numerical value:
"yes" (case insensitive) → 1
"no" → 0
"maybe" → 2
SciAuxThis repository contains the SciAux dataset, introduced in the paper Thinking in a Crowd: How Auxiliary Information Shapes LLM Reasoning.
SciAux is a new dataset derived from ScienceQA, designed to systematically test the robustness of Large Language Models (LLMs) against various types of auxiliary information (helpful, irrelevant, or misleading). The dataset aims to investigate the causal impact of such information on the reasoning process of LLMs with explicit step-by-step thinking… See the full description on the dataset page: https://huggingface.co/datasets/billhdzhao/SciAux.NCERT_Science_9thNCERT_Science_7thNCERT_Political_Science_11thscientific-agent-protocol-traces
SciAgentTrace
Matched cross-domain dataset of scientific-agent protocols. The central
comparison contains the same 6,653 problems under two actor models and four
protocols: 53,224 trajectories in 40 complete model--benchmark--protocol
groups. The broader table-first package contains 68,892
trajectories. Begin with trajectories, outcomes, or matched_outcomes, then
follow stable identifiers to messages and compressed raw traces.
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/scientific-agent-protocol-traces.sciknoweval-v2-hard-autogradable-512-2026-04-28
SciKnowEval v2 Hard Autogradable 512 - 2026-04-28
A 512-example sanity subset sampled from hicai-zju/SciKnowEval (v2, test) for Plan-CRL scientific reasoning evals.
Selection seed: 20260428.
Filtering and balancing:
excludes L1
keeps L2, L3, L4
keeps autogradable types: mcq-4-choices, mcq-2-choices, true_or_false, filling
requires answerKey or answer
balances domains at 128 examples each: Biology, Chemistry, Material, Physics
per domain: 32 L2, 48 L3, 48 L4
Useful fields for… See the full description on the dataset page: https://huggingface.co/datasets/1337xyz1337xyz/sciknoweval-v2-hard-autogradable-512-2026-04-28.SLM-in-SciPaper
SLM-in-SciPaper
This repository stores the data and model assets used by the SLM-in-SciPaper project.
Contents
data/keyword_keyphrase: processed Stage 1 keyphrase extraction data.
data/structure: processed Stage 2 structural evidence modeling data.
data/paper_corpus/full_library_txt: 178 plain-text scientific papers used as the local demonstration corpus.
data/paper_corpus/manifest.csv: metadata and file paths for the 178-paper text corpus.… See the full description on the dataset page: https://huggingface.co/datasets/KennySimpson/SLM-in-SciPaper.cleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research元データ: https://huggingface.co/datasets/moremilk/CoT_Reasoning_Scientific_Discovery_and_Research
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/CoT_Reasoning_Scientific_Discovery_and_Research
データ件数: 3,733
平均トークン数: 1,193
最大トークン数: 2,489
合計トークン数: 4,453,517
ファイル形式: JSONL
ファイル分割数: 1
合計ファイルサイズ: 23.2 MB
加工内容:
メタデータ列の解析と新列生成: metadata列(辞書型)を解析し、その中のreasoningをthought列に、difficultyをdifficulty列に展開しました。解析に失敗した行は除外されました。また、元のmetadata列は削除されました。
難易度によるフィルタリング:… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research.NCERT_Science_10th
