datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-SFT-Science-v2
Dataset Description:
Nemotron-Science-v2 is a science reasoning dataset with synthetic (synthetic MCQ, RQA) and non-synthetic vendor problems and LLM-generated solutions. It comprises three domains (Physics, Biology, and Chemistry), two question formats (multiple-choice questions [MCQ] and open questions [OpenQ]), and three generation setups: chain-of-thought (CoT) reasoning without tools, Python tool usage, and search tools usage with the Tavily API.
The solutions were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Science-v2.Nemotron-RL-Science-v1
Dataset Description:
Nemotron-RL-Science-v1 is a reinforcement learning (RL) dataset for science reasoning. Each example provides a problem, a reference answer, and a verifiable RL environment configuration (the agent prompt, the agent/verifier reference, and the answer-extraction template) so that a policy model can be trained with verifiable rewards. It covers three domains (Physics, Biology, and Chemistry), the open-question (OpenQ) format, and two generation setups:… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Science-v1.10001-Science-Facts
10,001 Science Facts
10,000+ obscure, surprising, and verifiable science facts
The kind that make you go "wait, really?"
🔗 GitHub Repository •
📁 Download by Category
🤔 What is this?
A curated dataset of 10,003 science facts across 32 categories — from quantum physics to parasites to the history of food.
Every fact is:
Sourced — from Wikipedia, Wikidata, academic sources
Verifiable — no LLM hallucinations
Surprising — passes the "dinner party test"… See the full description on the dataset page: https://huggingface.co/datasets/Royal-lobster/10001-Science-Facts.science-qa-sft-100k
Science QA SFT (100K)
100,000 science Q&A examples with step-by-step explanations for SFT fine-tuning. Covers physics, chemistry, biology, astronomy, and earth science at beginner through advanced difficulty.
Motivation
Models trained on general text often give superficially plausible but mechanistically wrong answers to science questions — stating the right conclusion without understanding the underlying reasoning. This dataset trains models to explain why an… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/science-qa-sft-100k.GLM-5.2-Science
GLM-5.2 · Science-50000x
50,000x traces distilled from GLM-5.2 on High reasoning
Physics · Chemistry · Biology
Token Count: 160M
Theres prompt overlap with my Kimi K2.5 dataset science subset, which I think those prompts are getting used in alot of places now
You can use this dataset for any purpose and you dont need to credit me, preferably dont claim it as your own.
hi - ianncity
10001-Science-Facts
10,001 Science Facts
10,000+ obscure, surprising, and verifiable science facts
The kind that make you go "wait, really?"
🔗 GitHub Repository •
📁 Download by Category
🤔 What is this?
A curated dataset of 10,003 science facts across 32 categories — from quantum physics to parasites to the history of food.
Every fact is:
Sourced — from Wikipedia, Wikidata, academic sources
Verifiable — no LLM hallucinations
Surprising — passes the "dinner party test"… See the full description on the dataset page: https://huggingface.co/datasets/percepteyeAI/10001-Science-Facts.sonnet3.5_science_conversationsThis dataset features sharegpt structured dialogues focused on a variety of advanced scientific topics. The content reflects a high level of scientific expertise, providing in-depth information on complex subjects.
Rehber-CoT-Science
🧬 Rehber-CoT-Science: Turkish Scientific Reasoning Dataset
Turkish Scientific Computational Reasoning (Chain-of-Thought) Dataset
Multi-step scientific problem-solving dataset with verifiable Python code and detailed explanations
Dataset • Author
📌 Changelog
Eski sürümlere erişim: Branch menüsünden v1 seçebilirsiniz.
Version
Date
Changes
v2.0
24.12.2025
✨ Yeni explained_answer alanı eklendi, Statistics domain eklendi, 712 örneğe genişletildi… See the full description on the dataset page: https://huggingface.co/datasets/batuhanozkose/Rehber-CoT-Science.nvidia-Nemotron-Science-Math
NVIDIA Nemotron Science and Math Reasoning
This is an unofficial, curated collection derived from NVIDIA's open-source Nemotron datasets. It is designed specifically to train language models in complex scientific and mathematical reasoning by providing structured chain-of-thought (CoT) examples.
To ensure efficiency, the shortest available CoT sequence was chosen for each question, filtering out redundant variations while preserving the core logical progression toward the final… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/nvidia-Nemotron-Science-Math.verisci-verified-science-math-code
VeriSci Verified Science Math Code
Verifier-grounded dataset for the Adaption AutoScientist Challenge Part 2, targeting the Science category with secondary Math and Code coverage.
Summary
VeriSci trains models to solve scientific computations, finite-difference PDE updates, numerical ODE steps, unit-checked mechanics, thermodynamics, circuits, chemistry stoichiometry, molarity, unit conversion, vector decomposition, two-point linear modeling, small Python… See the full description on the dataset page: https://huggingface.co/datasets/TheJackBright/verisci-verified-science-math-code.data-science-workflows-sft-100k
Data Science Workflows SFT (100K)
100,000 ShareGPT conversations demonstrating expert-level data science practice across data cleaning, EDA, ML pipelines, feature engineering, SQL analytics, statistical analysis, model evaluation, visualization, and production deployment.
Motivation
Data science is one of the most in-demand technical skills — companies need models that can reason through real analytical problems with the rigor of a senior data scientist. Models… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/data-science-workflows-sft-100k.deep-physics-science-zh
Deep Physics & Science Dialogue Dataset (Chinese)
深度物理科学对话数据集
Dataset Description
High-quality Chinese physics and science dialogues covering quantum gravity, theory of everything, relativity, quantum mechanics, and entropy/information theory.
高质量中文物理科学对话,涵盖量子引力理论、万物理论、相对论、量子力学、熵与信息论等硬核科学议题。
Dataset Structure
Format: JSONL (JSON Lines)
Fields:
instruction: User message / question
input: Additional context (if any)
output: AI response… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-physics-science-zh.science_behavioral_and_domain_diversity_dataset
Nepali Science SFT Dataset — Clean Candidate
A high-quality Nepali Science Supervised Fine-Tuning (SFT) dataset containing short question–answer instruction-following examples written primarily in Nepali Devanagari script.
This release is the clean candidate produced after structural validation, language checks, duplicate analysis, and Unicode-contamination filtering.
Dataset Overview
Property
Value
Dataset file
clean_candidate.jsonl
Records
29,320… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/science_behavioral_and_domain_diversity_dataset.BCE-Prettybird-Middle-Science-v0.1
BCE-Prettybird-Middle-Science-v0.1 - 101000 Science Q&A Dataset for Instruction-Based Learning
We are excited to introduce a comprehensive math-physics-chemistry-biology dataset containing 100500 instruction-based question-answer pairs, designed to support research in science reasoning, problem-solving, and AI training. Generated using Python’s math libraries (e.g., math, numpy, sympy), the dataset covers a diverse range of difficulty levels—from basic arithmetic and algebra to… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Middle-Science-v0.1.BCE-Prettybird-Nano-Science-v0.1
BCE-Prettybird-Nano-Science-v0.1 - 500 Science Q&A Dataset for Instruction-Based Learning
We are excited to introduce a comprehensive math-physics-chemistry-biology dataset containing 500 instruction-based question-answer pairs, designed to support research in science reasoning, problem-solving, and AI training. Generated using Python’s math libraries (e.g., math, numpy, sympy), the dataset covers a diverse range of difficulty levels—from basic arithmetic and algebra to advanced… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Science-v0.1.scbe-life-science-research-training-demo
Status: experimental. Experiment-specific slice. Primary public dataset: scbe-aethermoore-training-data.
SCBE Research Training Package
This package was generated from live pubmed pulls for the query protein structure prediction and is meant for
lightweight Hugging Face dataset and SFT experiments.
Files
papers.jsonl: normalized raw research records
sft_train.jsonl: train split for instruction-style tasks
sft_validation.jsonl: validation split… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-life-science-research-training-demo.ru_scienceThe dataset is based on russian scientific articles. Data filtering was easy, there may be garbage text
autoscientist-science-dataset
AutoScientist adapted dataset — science
Adaption Labs AutoScientist v5 adapted fine-tuning data for the science category.
science_adapted.jsonl — prompt/completion pairs used for QLoRA SFT.
science_v5_raw.csv — full Adaption output (prompt, completion, enhanced_prompt, chosen, rejected, reasoning_trace, embeddings) used for DPO.
Paired weights: Rishidar/autoscientist-science-qlora (Kaggle mirror rishidard/autoscientist-science-qlora).
data-science-chatbot
📊 Data Science Chatbot Dataset (2000 Samples)
🚀 A high-quality instruction-style dataset designed for fine-tuning Large Language Models (LLMs) on Data Science concepts.
This dataset contains ~2000 curated question-answer pairs in ChatML format, enabling models to learn how to explain, define, and discuss core data science topics in a clear and beginner-friendly way.
🎯 Objective
The goal of this dataset is to:
Train LLMs to act as a Data Science Tutor
Provide clear… See the full description on the dataset page: https://huggingface.co/datasets/Hamzasajjad38/data-science-chatbot.science-factcheck-indic
Health-Science Fact-Check (Hindi/Punjabi)
Native Hindi and Punjabi text from
ai4bharat/IndicCorpV2,
adapted with AutoScientist into substantive domain responses written by
a health-science fact-checker judging claims and showing the reasoning.
Rows
667
Unique source texts
667
Absolute quality score
9.0/10 (grade A)
Source score before adaptation
9.0/10 (grade A)
Percentile
33.0
Relative change
+0.0%
Median response length
520 chars… See the full description on the dataset page: https://huggingface.co/datasets/tojpaj/science-factcheck-indic.rus_science_for_gpt_oss_20b
rus_science_for_gpt_oss_20b
Русскоязычный датасет для дообучения LLM под научно-академический ассистент.
Описание
~32 272 примера в формате JSONL.
Тематика: научные тексты, академический стиль, описание таблиц/методик, введения, пояснения, переформулировки.
Каждая строка содержит полный контекст диалога и готовые ответы ассистента.
Формат полей
reasoning_language: язык рассуждений ("Russian").
developer: инструкция для ассистента (роль/стиль/задача).
user:… See the full description on the dataset page: https://huggingface.co/datasets/MIldoc/rus_science_for_gpt_oss_20b.grade3-science-explanations-v4r7
Grade-Level Science Explanations v4r7
The final 485-record supervised fine-tuning dataset for the grade-level science
explainer. Each record maps a unique elementary-science question to a concise,
mechanism-complete explanation. Training uses the minimal prompt
Explain: {phrasing} so the reading behavior must be learned from examples rather
than supplied through prompt instructions.
Files
File
Records
Purpose
gold_v4_r7.jsonl
485
Final training split… See the full description on the dataset page: https://huggingface.co/datasets/SAgarwal34/grade3-science-explanations-v4r7.
