datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
math-code-science-deepseek-r1-en
R1 Dataset Collection
Aggregated high-quality English prompts and model-generated responses from DeepSeek R1 and DeepSeek R1-0528.
Dataset Summary
The R1 Dataset Collection combines multiple public DeepSeek-generated instruction-response corpora into a single, cleaned, English-only JSONL file. Each example consists of a <|user|> prompt and a <|assistant|> response in one "text" field. This release includes:
~21,000 examples from the DeepSeek-R1-0528 Distilled Custom… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/math-code-science-deepseek-r1-en.BenchMAX_Science
Dataset Sources
Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Link: https://huggingface.co/papers/2502.07346
Repository: https://github.com/CONE-MT/BenchMAX
Dataset Description
BenchMAX_Science is a dataset of BenchMAX, sourcing from GPQA, which evaluates the natural science reasoning capability in multilingual scenarios.
We extend the original English dataset to 16 non-English languages.
The data is first translated by Google… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Science.chinese-materials-science-open-intelligence
🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset
Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.science_reasoning
science_reasoning
Mistral-7B의 과학 지식·추론 능력 향상을 위해 6개 공개 과학 객관식 QA 데이터셋을 통일 포맷으로 변환하고, ARC-Challenge test와의 오염을 제거한 데이터셋입니다.
원본 데이터셋
allenai/sciq
allenai/openbookqa (main)
allenai/qasc
allenai/quartz
allenai/ai2_arc (ARC-Easy / ARC-Challenge)
nguyen-brat/worldtree
전처리
포맷 통일: 각 데이터셋의 서로 다른 스키마를 unique_id, orig_id, source, question, choices, answer, support 필드로 변환. support는 근거 문단/문장으로, 데이터셋별 원본 필드(support/fact/para/cot)에서 구성하거나 없으면 빈 문자열.… See the full description on the dataset page: https://huggingface.co/datasets/seonjeongh/science_reasoning.10001-Science-Facts
10,001 Science Facts
10,000+ obscure, surprising, and verifiable science facts
The kind that make you go "wait, really?"
🔗 GitHub Repository •
📁 Download by Category
🤔 What is this?
A curated dataset of 10,003 science facts across 32 categories — from quantum physics to parasites to the history of food.
Every fact is:
Sourced — from Wikipedia, Wikidata, academic sources
Verifiable — no LLM hallucinations
Surprising — passes the "dinner party test"… See the full description on the dataset page: https://huggingface.co/datasets/Royal-lobster/10001-Science-Facts.science-qa-sft-100k
Science QA SFT (100K)
100,000 science Q&A examples with step-by-step explanations for SFT fine-tuning. Covers physics, chemistry, biology, astronomy, and earth science at beginner through advanced difficulty.
Motivation
Models trained on general text often give superficially plausible but mechanistically wrong answers to science questions — stating the right conclusion without understanding the underlying reasoning. This dataset trains models to explain why an… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/science-qa-sft-100k.GLM-5.2-Science
GLM-5.2 · Science-50000x
50,000x traces distilled from GLM-5.2 on High reasoning
Physics · Chemistry · Biology
Token Count: 160M
Theres prompt overlap with my Kimi K2.5 dataset science subset, which I think those prompts are getting used in alot of places now
You can use this dataset for any purpose and you dont need to credit me, preferably dont claim it as your own.
hi - ianncity
science-on-a-sphere-prompt-completions
Dataset Card for Science On a Sphere QA Dataset
Dataset Details
Dataset Description
This dataset comprises question-and-answer (QA) pairs generated from NOAA's Science On a Sphere (SOS) website, including support documentation and the dataset catalog. Each entry contains a prompt and a corresponding completion, designed to support educational and research use cases in Earth science.
This dataset includes a custom dataset_script.py and a consolidated file… See the full description on the dataset page: https://huggingface.co/datasets/HacksHaven/science-on-a-sphere-prompt-completions.VietEmbed-RAG-Science
VietEmbed-RAG Science
VietEmbed-RAG Science is a Vietnamese retrieval dataset containing 68,567 query-document examples across seven scientific and technical domains.
Each record consists of:
A Vietnamese query (anchor)
A relevant passage (positive)
A semantically related but non-answering passage (hard_negative)
Topic and domain metadata
The dataset is designed for training and domain adaptation of Vietnamese text embedding, semantic retrieval, and Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/nhminh107/VietEmbed-RAG-Science.food-science-llm-protocol
Food Science LLM Text-Mining Protocol
Pipeline and derived data accompanying:
Guo X, Fu W. Data Mining and Text Mining Using Large Language Models.
In: Li Y, Zhang D, Guo Z (eds), AI in Food Science: Methods and Protocols.
Methods and Protocols in Food Science. Springer.
The chapter prints one protocol as 26 numbered steps with abbreviated code
listings. This repository is the executable form of that protocol. Every step
has a corresponding function here, and every number in… See the full description on the dataset page: https://huggingface.co/datasets/KSU-HW-SEC/food-science-llm-protocol.10001-Science-Facts
10,001 Science Facts
10,000+ obscure, surprising, and verifiable science facts
The kind that make you go "wait, really?"
🔗 GitHub Repository •
📁 Download by Category
🤔 What is this?
A curated dataset of 10,003 science facts across 32 categories — from quantum physics to parasites to the history of food.
Every fact is:
Sourced — from Wikipedia, Wikidata, academic sources
Verifiable — no LLM hallucinations
Surprising — passes the "dinner party test"… See the full description on the dataset page: https://huggingface.co/datasets/percepteyeAI/10001-Science-Facts.verisci-verified-science-math-code
VeriSci Verified Science Math Code
Verifier-grounded dataset for the Adaption AutoScientist Challenge Part 2, targeting the Science category with secondary Math and Code coverage.
Summary
VeriSci trains models to solve scientific computations, finite-difference PDE updates, numerical ODE steps, unit-checked mechanics, thermodynamics, circuits, chemistry stoichiometry, molarity, unit conversion, vector decomposition, two-point linear modeling, small Python… See the full description on the dataset page: https://huggingface.co/datasets/TheJackBright/verisci-verified-science-math-code.Rehber-CoT-Science
🧬 Rehber-CoT-Science: Turkish Scientific Reasoning Dataset
Turkish Scientific Computational Reasoning (Chain-of-Thought) Dataset
Multi-step scientific problem-solving dataset with verifiable Python code and detailed explanations
Dataset • Author
📌 Changelog
Eski sürümlere erişim: Branch menüsünden v1 seçebilirsiniz.
Version
Date
Changes
v2.0
24.12.2025
✨ Yeni explained_answer alanı eklendi, Statistics domain eklendi, 712 örneğe genişletildi… See the full description on the dataset page: https://huggingface.co/datasets/batuhanozkose/Rehber-CoT-Science.nvidia-Nemotron-Science-Math
NVIDIA Nemotron Science and Math Reasoning
This is an unofficial, curated collection derived from NVIDIA's open-source Nemotron datasets. It is designed specifically to train language models in complex scientific and mathematical reasoning by providing structured chain-of-thought (CoT) examples.
To ensure efficiency, the shortest available CoT sequence was chosen for each question, filtering out redundant variations while preserving the core logical progression toward the final… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/nvidia-Nemotron-Science-Math.science_textbook_elementary_korBCE-Prettybird-Middle-Science-v0.1
BCE-Prettybird-Middle-Science-v0.1 - 101000 Science Q&A Dataset for Instruction-Based Learning
We are excited to introduce a comprehensive math-physics-chemistry-biology dataset containing 100500 instruction-based question-answer pairs, designed to support research in science reasoning, problem-solving, and AI training. Generated using Python’s math libraries (e.g., math, numpy, sympy), the dataset covers a diverse range of difficulty levels—from basic arithmetic and algebra to… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Middle-Science-v0.1.aalen_university_faculty_computer_science
Dataset Card
This dataset contains question-answer pairs from all study programmes of the Faculty of Computer Science at the University of Aalen, Germany. The training dataset is automatically generated by ChatGPT. The validation dataset was manually created.
It was collected to train an answer-Q&A chatbot based on LLM fine-tuning. All used scripts and examples can be found in the linked GitHub repository (https://github.com/pattplatt/llm_dataset_creation_and_finetuning).… See the full description on the dataset page: https://huggingface.co/datasets/Puidii/aalen_university_faculty_computer_science.nexa-science-multitask-balanced
Nexa Science Multitask Balanced
This dataset is a curated, instruction-formatted scientific multitask mixture for:
claim verification (<TASK:VERIFY>)
abstract-grounded biomedical QA (<TASK:QA>)
retrieval relevance re-ranking (<TASK:RERANK>)
Format
Each row is JSONL with:
{task, instruction, input, output, meta}
Splits Included
train_balanced_short.jsonl
val_balanced_short.jsonl
stats_balanced_short.json
Notes
QA in this balanced release is… See the full description on the dataset page: https://huggingface.co/datasets/AethronPhantom/nexa-science-multitask-balanced.science_behavioral_and_domain_diversity_dataset
Nepali Science SFT Dataset — Clean Candidate
A high-quality Nepali Science Supervised Fine-Tuning (SFT) dataset containing short question–answer instruction-following examples written primarily in Nepali Devanagari script.
This release is the clean candidate produced after structural validation, language checks, duplicate analysis, and Unicode-contamination filtering.
Dataset Overview
Property
Value
Dataset file
clean_candidate.jsonl
Records
29,320… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/science_behavioral_and_domain_diversity_dataset.BCE-Prettybird-Nano-Science-v0.1
BCE-Prettybird-Nano-Science-v0.1 - 500 Science Q&A Dataset for Instruction-Based Learning
We are excited to introduce a comprehensive math-physics-chemistry-biology dataset containing 500 instruction-based question-answer pairs, designed to support research in science reasoning, problem-solving, and AI training. Generated using Python’s math libraries (e.g., math, numpy, sympy), the dataset covers a diverse range of difficulty levels—from basic arithmetic and algebra to advanced… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Science-v0.1.scbe-life-science-research-training-demo
Status: experimental. Experiment-specific slice. Primary public dataset: scbe-aethermoore-training-data.
SCBE Research Training Package
This package was generated from live pubmed pulls for the query protein structure prediction and is meant for
lightweight Hugging Face dataset and SFT experiments.
Files
papers.jsonl: normalized raw research records
sft_train.jsonl: train split for instruction-style tasks
sft_validation.jsonl: validation split… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-life-science-research-training-demo.bro-science-benchmark
BroScienceBench
A literature-grounded benchmark for evaluating how large language models handle
strength-training misinformation ("bro-science"). 246 items across 42 myth
clusters and 7 categories, each pairing an evidence-based answer against a
documented gym myth and a plausible distractor.
Code, evaluation harness, analysis scripts, figures, and the full datasheet:
https://github.com/puranjayh/bro-science-benchmark
⚠️ For engineering evaluation and research only. Not medical… See the full description on the dataset page: https://huggingface.co/datasets/pur4nj41y/bro-science-benchmark.science-sinhala-gce-olevel-2023-mcq
Dataset Details
This dataset contains 40 Science MCQ questions and answers in Sinhala language of the GCE Ordinary Level Science paper 2023.
ScienceQA-Weather-R1
Introduction
Dataset Summary
This dataset is the out-of-domain (OOD) evaluation set used in "Weather-R1: Logically Consistent Reinforcement Fine-Tuning for Multimodal Reasoning in Meteorology". It is curated from the "Weather and climate" category of ScienceQA and contains 324 English, multimodal multiple-choice questions to test cross-domain generalization.
Supported Tasks
Multi-modal Multiple Choice
Languages
English
Dataset Overview
For… See the full description on the dataset page: https://huggingface.co/datasets/Marco711/ScienceQA-Weather-R1.science-qa-samples
Science Q&A Samples
This sample shows structured science question-answer pairs for reviewing subject coverage, difficulty labeling, and answer format before scoping a larger educational dataset.
What This Shows
Q&A examples across science and math subjects
Metadata for topic, difficulty, curriculum alignment, and question type
A view of how text and asset-backed questions are represented
Dataset Specifications
Field
Value
Modality
Text… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/science-qa-samples.science_textbook_elementary_kor_seedsinhala-political-science-gce-alevel-2021-questionsdata-science-chatbot
📊 Data Science Chatbot Dataset (2000 Samples)
🚀 A high-quality instruction-style dataset designed for fine-tuning Large Language Models (LLMs) on Data Science concepts.
This dataset contains ~2000 curated question-answer pairs in ChatML format, enabling models to learn how to explain, define, and discuss core data science topics in a clear and beginner-friendly way.
🎯 Objective
The goal of this dataset is to:
Train LLMs to act as a Data Science Tutor
Provide clear… See the full description on the dataset page: https://huggingface.co/datasets/Hamzasajjad38/data-science-chatbot.
