datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
China-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.AI-Research-Evaluation-Repository-STEM
AI-STEM-Research-Eval-Dataset
Overview
This dataset contains AI-generated scientific reports across STEM domains, accompanied by structured metadata, prompt documentation, reference validation, and hallucination annotations.
It is designed as an open research resource to study the capabilities, limitations, and reliability of large language models (LLMs) in generating scientific content.
The dataset enables systematic analysis of how AI systems perform in… See the full description on the dataset page: https://huggingface.co/datasets/sreearravind/AI-Research-Evaluation-Repository-STEM.stem-reasoning-complex
STEM-Reasoning-Complex: High-Fidelity Scientific CoT Dataset
1. Dataset Summary
STEM-Reasoning-Complex is a curated collection of 118.255 high-quality samples designed for Supervised Fine-Tuning (SFT) and alignment of Large Language Models. The dataset focuses on four core disciplines: Biology, Mathematics, Physics, and Chemistry.
Unlike standard QA datasets, each entry provides a structured Chain-of-Thought (CoT) reasoning process, enabling models to learn… See the full description on the dataset page: https://huggingface.co/datasets/galaxyMindAiLabs/stem-reasoning-complex.shreyansh-1B-SLM-pretrain-stem-english
📚 Vigyan Pretrain Corpus: 13GB Web-Scale Scientific & Technical Text
The Vigyan Pretrain Corpus is a web-scale, curated raw text pre-training dataset comprising 13.14 GB of high-density STEM literature, textbooks, open-access research papers, and technical documentations.
🔬 Dataset Overview
Designed specifically for pre-training and continuous pre-training (CPT) of Small Language Models (SLMs) in the 1B–3B parameter regime:
High Information Density: Filtered to… See the full description on the dataset page: https://huggingface.co/datasets/shreyansh12183/shreyansh-1B-SLM-pretrain-stem-english.stem-scientific-code-sample
AxiomSet Labs STEM Scientific-Code Sample
A 30-task sample of STEM reasoning and scientific-code problems across five domains.
Domains
Biology: 6 tasks
Chemistry: 6 tasks
Materials Science: 6 tasks
Mathematics: 6 tasks
Physics: 6 tasks
Each task contains two subproblems and one main problem, including prompts, scientific background, testing templates, and reference solutions.
Files
data/sample.jsonl — one task per line; used by the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/AxiomSetLabs/stem-scientific-code-sample.tushe-grade-school-stem
Tushe Community Grade School STEM
Open dataset of grade-school STEM (Science, Technology, Engineering, Mathematics) textbooks, curated for Tushe Community and aligned with curriculum use (e.g. CAPS-aligned content).
Data Fields (per book JSON)
Field
Type
Description
source_file
string
Original .txt filename
title
string
Derived book title (e.g. "Grade 8A Mathematics")
table_of_contents
list
[{ "section_id", "title" }, ...]
front_matter
string
Intro… See the full description on the dataset page: https://huggingface.co/datasets/Tushe/tushe-grade-school-stem.gpt-oss-120b-reasoning-STEM-5K
GPT-OSS-120B-Distilled-Reasoning-STEM Dataset
1) Dataset Overview
Data Source Model: gpt-oss-120b-high
Task Type: STEM Reasoning and Problem Solving (Science, Technology, Engineering & Mathematics)
Data Format: `JSON Lines
Fields: generator, category, input, CoT_Native——reasoning, answer
(Consistent with the math dataset, splitting the original 'output' into 'reasoning' and 'answer' for COT/SFT scenarios.)
2) Design Goals (Motivation)
This dataset targets… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120b-reasoning-STEM-5K.salabs-stem-deep-reasoning-cot-v13
🧪 SALabs Multi-Domain STEM Deep Reasoning & Chain-of-Thought (CoT) Corpus (v13.0)
[!IMPORTANT]
💳 Click Here to Purchase Enterprise Commercial License ($2,500 USD) & Instant 31.7MB Master Archive DownloadInstant download of the full lossless master package containing all 1,816 JSONL reasoning records + 13 complete uncompressed text corpora (31.72 MB uncompressed total) + commercial license certificate.
🌟 Executive Summary
The SALabs STEM Deep Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/suitai/salabs-stem-deep-reasoning-cot-v13.Electrical-engineering
To the electrical engineering community
This dataset contains Q&A prompts about electrical engineering, Kicad's EDA software features and scripting console Python codes.
Authors
STEM.AI: stem.ai.mtl@gmail.comWilliam Harbec
stem-reasoning-v1.0.0-ccbysa-001
YouAI Data — stem-reasoning-v1.0.0-ccbysa-001
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 394 step-by-step reasoning chains and 569 instruction/response pairs across 332 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 364 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-v1.0.0-ccbysa-001.stem-reasoning-ccbysa-006
YouAI Data — stem-reasoning-ccbysa-006
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 251 step-by-step reasoning chains and 729 instruction/response pairs across 662 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 266 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-ccbysa-006.stem-reasoning-ccbysa-002
YouAI Data — stem-reasoning-ccbysa-002
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 529 step-by-step reasoning chains and 439 instruction/response pairs across 591 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 408 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-ccbysa-002.stem-tr-instruct-1k
Eding STEM TR Instruct 1K
Türkçe K-12 STEM ve kodlama eğitimi için instruction-tuning veri seti.
Veri Seti Hakkında
Bu veri seti, Türkiye'deki K-12 seviyesinde STEM ve kodlama eğitimi için
hazırlanmış 1.000 instruction-output çiftinden oluşur.
Kategoriler
Arduino: LED, sensör, motor projeleri, devre tasarımı
Scratch: Blok tabanlı programlama, oyun yapımı, animasyon
mBlock: mBot robot programlama, sensör kullanımı
Robotik: PID kontrol, çizgi izleme… See the full description on the dataset page: https://huggingface.co/datasets/alimkacar/stem-tr-instruct-1k.stem-reasoning-ccbysa-008
YouAI Data — stem-reasoning-ccbysa-008
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 73 step-by-step reasoning chains and 921 instruction/response pairs across 667 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 361 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-ccbysa-008.shreyansh-hinglish-english-stem-500k
🇮🇳 Vigyan Indic-STEM: 500k Bilingual Hinglish & English Reasoning Corpus
Vigyan Indic-STEM 500k is a specialized, large-scale bilingual dataset created to bridge the pedagogical divide in STEM education across India. It pairs rigorous English first-principles scientific derivations with natural, conversational Hinglish (Hindi written in Roman script) explanations.
📖 Overview
In Tier-2 and Tier-3 educational institutions across India, STEM concepts (Physics… See the full description on the dataset page: https://huggingface.co/datasets/shreyansh12183/shreyansh-hinglish-english-stem-500k.stem-reasoning-ccbysa-011
YouAI Data — stem-reasoning-ccbysa-011
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 11 step-by-step reasoning chains and 989 instruction/response pairs across 454 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 563 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-ccbysa-011.stem-reasoning-ccbysa-005
YouAI Data — stem-reasoning-ccbysa-005
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 248 step-by-step reasoning chains and 729 instruction/response pairs across 696 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 559 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-ccbysa-005.stem-reasoning-ccbysa-001
YouAI Data — stem-reasoning-ccbysa-001
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 277 step-by-step reasoning chains and 719 instruction/response pairs across 623 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 584 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-ccbysa-001.stem-reasoning-complex
STEM-Reasoning-Complex: High-Fidelity Scientific CoT Dataset
1. Dataset Summary
STEM-Reasoning-Complex is a curated collection of 118.225 high-quality samples designed for Supervised Fine-Tuning (SFT) and alignment of Large Language Models. The dataset focuses on four core disciplines: Biology, Mathematics, Physics, and Chemistry.
Unlike standard QA datasets, each entry provides a structured Chain-of-Thought (CoT) reasoning process, enabling models to learn "internal… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/stem-reasoning-complex.stem-reasoning-ccbysa-007
YouAI Data — stem-reasoning-ccbysa-007
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 21 step-by-step reasoning chains and 979 instruction/response pairs across 674 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 585 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-ccbysa-007.quadmix-stem-v1
QuaDMix-STEM v1: STEM-Focused Proxy Validation Set
Script: scripts/validation_set/prepare_stem_v1.py
HuggingFace: liujin99/quadmix-stem-v1
Files: stem_v1_tokenized.pt, stem_v1.parquet
Overview
STEM v1 is a validation set designed to focus the proxy model's optimization signal on STEM capabilities — mathematics, science knowledge, and logical reasoning. Unlike CAP v1 (broad capability coverage) or core_bmk (benchmark test format), STEM v1 uses only tasks that… See the full description on the dataset page: https://huggingface.co/datasets/liujin99/quadmix-stem-v1.stem-reasoning-ccbysa-10k-001
stem-reasoning-ccbysa-10k-001
Combined 10K CC-BY-SA STEM reasoning dataset — a sequential merge of ten 1K sets.
Total examples: 10000
License: CC-BY-SA-4.0
Mean quality score: 4.411 (min 4.25, max 4.81)
Unique sources: 2387
DPO pairs: 3650
Difficulty: {'advanced': 2774, 'introductory': 5409, 'intermediate': 1817}
Products: {'reasoning_chain': 1531, 'instruction_pair': 8377, 'code_instruction': 92}
Combined from
stem-reasoning-ccbysa-001… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-ccbysa-10k-001.stem-reasoning-ccbysa-012
YouAI Data — stem-reasoning-ccbysa-012
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 72 step-by-step reasoning chains and 928 instruction/response pairs across 474 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 531 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-ccbysa-012.omni-stem
GitHub
Website
Paper (Coming Soon)
Dataset Details
This dataset is a combination of all datasets used for expert finetuning in a curriculum order of Math, Science, Technology, and Engineering based on similarities between each domain’s fundamentals, aiding model learning. This dataset contains a mixture of chat-based and continued-pretrain data.
Sources
This dataset was sourced from the following open-sourced datasets:
Math
meta-math/MetaMathQA… See the full description on the dataset page: https://huggingface.co/datasets/omniomni/omni-stem.reasoning-sft-stem-reasoning-complex-FineProofs-126K
reasoning-sft-stem-reasoning-complex-FineProofs-126K
Combined converted dataset from two sources:
lm-provers/FineProofs-SFT (all config, 7.78k) — Mathematical Olympiad problems with chain-of-thought reasoning distilled from DeepSeek-Math-V2
galaxyMindAiLabs/stem-reasoning-complex (~118k) — STEM reasoning across Biology, Mathematics, Physics, Chemistry and Code
Format
Each row has three columns:
input — list of dicts [{"role": "user", "content": "..."}]
response —… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-stem-reasoning-complex-FineProofs-126K.stem-reasoning-ccbysa-004
YouAI Data — stem-reasoning-ccbysa-004
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 22 step-by-step reasoning chains and 978 instruction/response pairs across 577 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 84 DPO preference pairs as a free companion dataset.… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-ccbysa-004.hausa-stem-reasoning-with-cultural-context
Hausa STEM Reasoning with Cultural Context
Abstract
We present the first large-scale bilingual Hausa-English STEM reasoning dataset with deep cultural adaptation, containing 2,640 high-quality question-answer pairs translated from the STEM-Reasoning-Complex dataset. Our work introduces the "Shehin Malamin Kimiyya" (The Wise Scholar of Science) translation framework, which transforms Western scientific concepts into culturally-embedded Hausa explanations using systematic… See the full description on the dataset page: https://huggingface.co/datasets/Tushe/hausa-stem-reasoning-with-cultural-context.quadmix-stem-v2
QuaDMix-STEM v2: STEM-Focused Proxy Validation Set with GPQA & MATH
Script: scripts/validation_set/prepare_stem_v2.py
HuggingFace: liujin99/quadmix-stem-v2
Files: stem_v2_tokenized.pt, stem_v2.parquet
Overview
STEM v2 is an upgraded validation set that fixes the two critical coverage gaps in STEM v1. In the v1 experiment, QuaDMix lost to Random downstream (CORE 0.1530 vs 0.1615), and root-cause analysis revealed:
gpqa_diamond had no direct proxy — mapped from… See the full description on the dataset page: https://huggingface.co/datasets/liujin99/quadmix-stem-v2.stem-reasoning-ccbysa-010
YouAI Data — stem-reasoning-ccbysa-010
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 7 step-by-step reasoning chains and 991 instruction/response pairs across 325 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 520 DPO preference pairs as a free companion dataset.… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-ccbysa-010.nanochat-npu-stem-eval
nanochat-npu-stem-eval
Pre-processed STEM evaluation data for nanochat-npu, adapted from karpathy/nanochat for Huawei 910B3 NPU.
Tasks
Task
Type
Shot
Source
Examples
Description
gpqa_diamond
multiple_choice
0-shot
Idavidrein/gpqa
198
Graduate-level science QA (Diamond subset)
gsm8k_cot
generation
8-shot
openai/gsm8k
1319
Grade school math word problems (CoT)
math_cot
generation
4-shot
HuggingFaceH4/MATH-500
500
Competition mathematics (CoT)… See the full description on the dataset page: https://huggingface.co/datasets/Sexhuis/nanochat-npu-stem-eval.
