datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dclm-stem-filteredeai-taxonomy-stem-w-dclm
🔬 EAI-Taxonomy STEM w/ DCLM
🏆 Website | 🖥️ Code | 📖 Paper
A high-quality STEM dataset curated from web data using taxonomy-based filtering, containing 1742 billion tokens of science, technology, engineering, and mathematics content.
🎯 Dataset Overview
This dataset is part of the Essential-Web project, which introduces a new paradigm for dataset curation using expressive metadata and simple semantic filters. Unlike traditional STEM datasets that require… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/eai-taxonomy-stem-w-dclm.arabic-stem-lexicon
Arabic Diacritized-Stem Lexicon
An undiacritized Arabic surface form → its most frequent diacritized stem.
Standard Arabic writes no short vowels, so anything that has to pronounce Arabic
must first put them back. A neural diacritizer does that well on rare words, where
inference is the only thing there is. On common words it is the wrong tool:
which vowels كتاب carries is not a thing to be inferred, it is a thing to be looked
up — and models get exactly these wrong, reading… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-stem-lexicon.MMLU-STEMThis contains a subset of STEM subjects defined in MMLU by the original paper.
The included subjects are
'abstract_algebra',
'anatomy',
'astronomy',
'college_biology',
'college_chemistry',
'college_computer_science',
'college_mathematics',
'college_physics',
'computer_security',
'conceptual_physics',
'electrical_engineering',
'elementary_mathematics',
'high_school_biology',
'high_school_chemistry',
'high_school_computer_science',
'high_school_mathematics',
'high_school_physics'… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/MMLU-STEM.melodix-stemsSTEM
STEM Dataset
📃 [Paper] • 💻 [Github] • 🤗 [Dataset] • 🏆 [Leaderboard] • 📽 [Slides] • 📋 [Poster]
This dataset is proposed in the ICLR 2024 paper: Measuring Vision-Language STEM Skills of Neural Models. We introduce a new challenge to test the STEM skills of neural models. The problems in the real world often require solutions, combining knowledge from STEM (science, technology, engineering, and math). Unlike existing datasets, our dataset requires the understanding of… See the full description on the dataset page: https://huggingface.co/datasets/stemdataset/STEM.eai-taxonomy-stem-w-dclm-100b-sample
🔬 EAI-Taxonomy STEM w/ DCLM (100B sample)
🏆 Website | 🖥️ Code | 📖 Paper
A high-quality STEM dataset curated from web data using taxonomy-based filtering, containing 100 billion tokens of science, technology, engineering, and mathematics content.
🎯 Dataset Overview
This dataset is part of the Essential-Web project, which introduces a new paradigm for dataset curation using expressive metadata and simple semantic filters. Unlike traditional STEM datasets that… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/eai-taxonomy-stem-w-dclm-100b-sample.Logics-STEM-SFT-Dataset-Open-1.6M
Logics-STEM-SFT-Dataset-2.2M
📰 News
[2026.01.05]🔥 Release of our Techinical Report.
[2026.01.05]🔥 Release the first version of Logics-STEM-8B-SFT, Logics-STEM-8B-RL, /Logics-STEM-SFT-Dataset-Open-1.6M.
Overview
What is this dataset?
Logics-STEM-SFT-Dataset-2.2M is a curated long Chain-of-Thought (CoT) SFT dataset for STEM reasoning, built on top of high-quality open-source data and enhanced through a rigorous curation and distillation… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-STEM-SFT-Dataset-Open-1.6M.stemdata
STEM Dataset
📃 [Paper] • 💻 [Github] • 🤗 [Dataset] • 🏆 [Leaderboard] • 📽 [Slides] • 📋 [Poster]
This dataset is proposed in the ICLR 2024 paper: Measuring Vision-Language STEM Skills of Neural Models. We introduce a new challenge to test the STEM skills of neural models. The problems in the real world often require solutions, combining knowledge from STEM (science, technology, engineering, and math). Unlike existing datasets, our dataset requires the understanding of… See the full description on the dataset page: https://huggingface.co/datasets/si-m07/stemdata.STEM2Crystal-Bench
STEM2Crystal-Bench
STEM2Crystal-Bench is the benchmark for the paper "From Noisy STEM to Crystal Structure: Evidence-Structure CoDiffusion under Composition Constraints" (Chen & You, KDD 2026, Oral), which introduces STEM2Crystal CoDiffusion (SCCD). It evaluates methods that reconstruct a crystal structure from a noisy STEM image when the composition is known. The release has a large synthetic set with controlled noise and a small set of real STEM images, with ground-truth CIFs… See the full description on the dataset page: https://huggingface.co/datasets/gary23ai/STEM2Crystal-Bench.stem-corpusChina-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.AI-Research-Evaluation-Repository-STEM
AI-STEM-Research-Eval-Dataset
Overview
This dataset contains AI-generated scientific reports across STEM domains, accompanied by structured metadata, prompt documentation, reference validation, and hallucination annotations.
It is designed as an open research resource to study the capabilities, limitations, and reliability of large language models (LLMs) in generating scientific content.
The dataset enables systematic analysis of how AI systems perform in… See the full description on the dataset page: https://huggingface.co/datasets/sreearravind/AI-Research-Evaluation-Repository-STEM.stem-reasoning-complex
STEM-Reasoning-Complex: High-Fidelity Scientific CoT Dataset
1. Dataset Summary
STEM-Reasoning-Complex is a curated collection of 118.255 high-quality samples designed for Supervised Fine-Tuning (SFT) and alignment of Large Language Models. The dataset focuses on four core disciplines: Biology, Mathematics, Physics, and Chemistry.
Unlike standard QA datasets, each entry provides a structured Chain-of-Thought (CoT) reasoning process, enabling models to learn… See the full description on the dataset page: https://huggingface.co/datasets/galaxyMindAiLabs/stem-reasoning-complex.Natural-Reasoning-STEM-25Kpat-stem-pretrainpfnc-gst-haadf-stem-eds-tomography-b2-d3
PFNC GST HAADF-STEM/EDS Tomography — B2 VIRGIN and D3 SET
Summary
This release contains two limited-angle HAADF-STEM/EDS tomography acquisitions of cross-sectional Ge-Sb-Te phase-change-memory devices acquired at the Platform for Nanocharacterisation (PFNC), CEA Grenoble, on a Thermo Fisher Scientific Titan Themis operated at 200 kV with a four-detector Super-X EDS system.
B2 and D3 are internal CEA microscope sample identifiers. In this release, B2 is the VIRGIN… See the full description on the dataset page: https://huggingface.co/datasets/scitomo/pfnc-gst-haadf-stem-eds-tomography-b2-d3.STEM2Mat
AutoMat Benchmark: STEM Image to Crystal Structure
The AutoMat Benchmark is a multimodal dataset designed to evaluate deep‑learning systems for iDPC-STEM‑based crystal‑structure reconstruction and property prediction.
Code: https://github.com/yyt-2378/AutoMat
📁 Dataset Structure
The dataset is organized into three tiers of increasing difficulty:
benchmark/
├── tier1/
│ ├── img/ # STEM images (e.g., PNG, TIFF)
│ ├── label/ # Atomic position labels… See the full description on the dataset page: https://huggingface.co/datasets/yaotianvector/STEM2Mat.shreyansh-1B-SLM-pretrain-stem-english
📚 Vigyan Pretrain Corpus: 13GB Web-Scale Scientific & Technical Text
The Vigyan Pretrain Corpus is a web-scale, curated raw text pre-training dataset comprising 13.14 GB of high-density STEM literature, textbooks, open-access research papers, and technical documentations.
🔬 Dataset Overview
Designed specifically for pre-training and continuous pre-training (CPT) of Small Language Models (SLMs) in the 1B–3B parameter regime:
High Information Density: Filtered to… See the full description on the dataset page: https://huggingface.co/datasets/shreyansh12183/shreyansh-1B-SLM-pretrain-stem-english.ultrafine-stem-part-2tushe-grade-school-stem
Tushe Community Grade School STEM
Open dataset of grade-school STEM (Science, Technology, Engineering, Mathematics) textbooks, curated for Tushe Community and aligned with curriculum use (e.g. CAPS-aligned content).
Data Fields (per book JSON)
Field
Type
Description
source_file
string
Original .txt filename
title
string
Derived book title (e.g. "Grade 8A Mathematics")
table_of_contents
list
[{ "section_id", "title" }, ...]
front_matter
string
Intro… See the full description on the dataset page: https://huggingface.co/datasets/Tushe/tushe-grade-school-stem.stem-scientific-code-sample
AxiomSet Labs STEM Scientific-Code Sample
A 30-task sample of STEM reasoning and scientific-code problems across five domains.
Domains
Biology: 6 tasks
Chemistry: 6 tasks
Materials Science: 6 tasks
Mathematics: 6 tasks
Physics: 6 tasks
Each task contains two subproblems and one main problem, including prompts, scientific background, testing templates, and reference solutions.
Files
data/sample.jsonl — one task per line; used by the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/AxiomSetLabs/stem-scientific-code-sample.STEM_sftLogics-STEM-SFT-Dataset-Open-5.3MCeltic_Stems_Reference_Sessions_Preview
Harmonic Frontier Audio – Celtic Constellation Reference Sessions (Preview, v0.9)
A high-fidelity music-production dataset designed to connect isolated source performances, production processing, arrangement context, and finished musical outcomes.
Celtic Constellation Reference Sessions (Preview), created by Harmonic Frontier Audio, introduces the Reference Sessions product vertical through a compact proof-of-concept built around purpose-recorded Celtic ensemble material.… See the full description on the dataset page: https://huggingface.co/datasets/Harmonic-Frontier-Audio/Celtic_Stems_Reference_Sessions_Preview.ultrafine-stem-part-1gpt-oss-120b-reasoning-STEM-5K
GPT-OSS-120B-Distilled-Reasoning-STEM Dataset
1) Dataset Overview
Data Source Model: gpt-oss-120b-high
Task Type: STEM Reasoning and Problem Solving (Science, Technology, Engineering & Mathematics)
Data Format: `JSON Lines
Fields: generator, category, input, CoT_Native——reasoning, answer
(Consistent with the math dataset, splitting the original 'output' into 'reasoning' and 'answer' for COT/SFT scenarios.)
2) Design Goals (Motivation)
This dataset targets… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120b-reasoning-STEM-5K.stem-diagrams
STEM Diagrams
30,325 technical diagrams (block diagrams, schematics, flowcharts, architectures)
extracted from arXiv papers across six engineering fields, each with a source
attribution and a quality score. Built by an LLM-curated pipeline and used to show
that a small frozen-feature classifier can replace the paid LLM labeling gate.
Paper: Distilling an LLM Diagram-Curation Pipeline into Local Classifiers (Adnan Abbasi, Thothica, 2026)
Code:… See the full description on the dataset page: https://huggingface.co/datasets/aeyxen/stem-diagrams.swahili-speech
Swahili Speech-to-Text Dataset
This dataset contains paired audio and text data for training and evaluating speech-to-text models in Swahili. The audio files have been processed to remove silence, converted to 44.1kHz mono FLAC format, and are paired with corresponding transcriptions.
Structure
audio_*.flac: Audio files in FLAC format, named by their corresponding text corpus ID.
metadata.jsonl: JSON Lines file with metadata for each audio-text pair. Each line is a JSON… See the full description on the dataset page: https://huggingface.co/datasets/stem-content-ai-project/swahili-speech.toksuite_stem
Dataset Card for Tokenization Robustness
TokSuite Benchmark (STEM Collection)
Dataset Description
This dataset is the STEM subset of the TokSuite benchmark, designed to evaluate how tokenizer choice affects model behavior under realistic formatting, notation, and surface-form perturbations in technical text. TokSuite includes specialized benchmarks for mathematics and STEM, with the STEM subset containing 44 canonical technical questions paired with a… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_stem.
