datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scientific-verification
Scientific Verification Benchmark: NMC Cathodes
Dataset summary
The benchmark contains 50 scientific claims about NMC (lithium nickel manganese cobalt oxide) battery cathodes. Each claim is answered by Claude Opus 5, GPT 5.6 Luna and Gemini 3.1 Pro using a set of 20 open-access papers, producing 150 scored answers. The accompanying reference set contains 1,991 experiment-grounded measurements curated from 227 open-access papers, with experimental conditions and… See the full description on the dataset page: https://huggingface.co/datasets/GenData-Research/scientific-verification.SciPredict
SciPredict: Can LLMs Predict the Outcomes of Research Experiments?
Paper: SciPredict: Can LLMs Predict the Outcomes of Research Experiments in Natural Sciences?
Overview
SciPredict is a benchmark evaluating whether AI systems can predict experimental outcomes in physics, biology, and chemistry. The dataset comprises 405 questions derived from recently published empirical studies (post-March 2025), spanning 33 subdomains.
Dataset Structure
Total Questions: 405… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/SciPredict.sciqa-thinking
sciqa-thinking
Randomly extracted 3000 rows from sciq and prompting Qwen3-14b to generate the intermediate reasoning traces, we created this dataset.
This should be used for LLM post-training, especially RL.
nasa-science-repos-sme-benchmark
NASA Science Repos SME Benchmark
A benchmark dataset for evaluating retrieval systems on NASA science repository discovery tasks. This dataset contains expert queries, a corpus of NASA science GitHub repositories, and relevance judgments.
Dataset Structure
Files
├── corpus.jsonl # 5,264 repositories with full metadata
├── queries.jsonl # 219 expert queries
└── qrels/
├── earth.tsv # Earth Science relevance judgments (162)
├──… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-repos-sme-benchmark.Dense-Information-Science-Physics-Dataset
Dense Information With Multiple Fine-tuned Variations
This dataaset has multiple for each input to learn how to express the same answer in different ways
Dataset Structure
The dataset contains two columns:
Column
Description
input
A science or quantum-physics question
output
A conversational answer to the question
Example:
{
"input": "What is quantum entanglement?",
"output": "Quantum entanglement is when two quantum systems share one… See the full description on the dataset page: https://huggingface.co/datasets/StarpowerTechnology/Dense-Information-Science-Physics-Dataset.scidcc-instructions
Dataset Summary
Instruction-Response pairs generated using the SciDCC Climate Dataset from Climabench
Format
### Instruction:
Present a fitting title for the provided text.
For those who study earthquakes, one major challenge has been trying to understand all the physics of a fault -- both during an earthquake and at times of "rest" -- in order to know more about how a particular region may behave in the future. Now, researchers at the California Institute of Technology… See the full description on the dataset page: https://huggingface.co/datasets/tanmaylaud/scidcc-instructions.ScienceOlympiad.tsv
Dataset Card for ScienceOlympiad.tsv
ScienceOlympiad.tsv: Challenging AI with Olympiad-Level Multimodal Science Problems.
Source: https://huggingface.co/datasets/ByteDance-Seed/ScienceOlympiad
Dataset Details
Dataset Description
The ScienceOlympiad dataset is a meticulously curated benchmark designed to evaluate the scientific reasoning capabilities of state-of-the-art AI models. It features elite, competition-level problems in physics and chemistry, addressing… See the full description on the dataset page: https://huggingface.co/datasets/YuJJJJin/ScienceOlympiad.tsv.Medi-Science
Medi-Science Dataset
The Medi-Science dataset is a comprehensive collection of medical Q&A data designed for text generation, question answering, and summarization tasks in the healthcare domain.
Dataset Overview
Name: Medi-Science
License: Apache-2.0
Languages: English
Tags: Medical, Medicine, Anomaly, Biology, Medi-Science
Number of Rows: 16,412
Dataset Size:
Downloaded: 22.7 MB
Auto-converted Parquet: 8.94 MB
Dataset Structure
The dataset contains the… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Medi-Science.NCERT_Science_10thscience_leaderboard_submissionThis dataset contains the results used for Science Leaderboard
NCERT_Science_8thNCERT_Political_Science_12thscips_qascikit-learn-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: scikit-learn
Data Source Link: https://scikit-learn.org/stable/index.html
Data Source License: https://github.com/scikit-learn/scikit-learn/blob/main/COPYING
Data Source Authors: scikit-learn contributors
AI Benchmarks by Data Agents. 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai
NCERT_Science_6thNCERT_Science_9thscipy-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: scipy documentation
Data Source Link: https://docs.scipy.org/doc/scipy/reference/index.html
Data Source License: CC0 1.0 Universal
Data Source Authors: https://docs.scipy.org/doc/scipy/dev/governance.html
AI Benchmarks by Data Agents. 2025 RELAI.AI Licensed under CC BY 4.0. Source: https://relai.ai
NCERT_Science_7thNCERT_Political_Science_11theveryday-science
Everyday Science Dataset
Dataset Summary
The "Everyday Science" dataset is a synthetically generated collection of question-answer pairs covering a diverse range of topics within everyday science. Each entry includes a topic, a specific question, a scientific explanation or principle, a relatable everyday example, relevant scientific concepts (keywords), and further reasoning or inference.
This dataset contains question-and-answer pairs focused on everyday science… See the full description on the dataset page: https://huggingface.co/datasets/nirajandhakal/everyday-science.scipy-standardSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: scipy documentation
Data Source Link: https://docs.scipy.org/doc/scipy/reference/index.html
Data Source License: CC0 1.0 Universal
Data Source Authors: https://docs.scipy.org/doc/scipy/dev/governance.html
AI Benchmarks by Data Agents. 2025 RELAI.AI Licensed under CC BY 4.0. Source: https://relai.ai
NCERT_Science_10thSFT_Science_AI_Genscikit-learn-standardSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: scikit-learn
Data Source Link: https://scikit-learn.org/stable/index.html
Data Source License: https://github.com/scikit-learn/scikit-learn/blob/main/COPYING
Data Source Authors: scikit-learn contributors
AI Benchmarks by Data Agents. 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai
clinical_sciences_dataThis is a dataset for training AI on medical tools and practices in the modern age.
