datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dclm-stem-filteredLogics-STEM-SFT-Dataset-Open-1.6M
Logics-STEM-SFT-Dataset-2.2M
📰 News
[2026.01.05]🔥 Release of our Techinical Report.
[2026.01.05]🔥 Release the first version of Logics-STEM-8B-SFT, Logics-STEM-8B-RL, /Logics-STEM-SFT-Dataset-Open-1.6M.
Overview
What is this dataset?
Logics-STEM-SFT-Dataset-2.2M is a curated long Chain-of-Thought (CoT) SFT dataset for STEM reasoning, built on top of high-quality open-source data and enhanced through a rigorous curation and distillation… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-STEM-SFT-Dataset-Open-1.6M.STEM2Crystal-Bench
STEM2Crystal-Bench
STEM2Crystal-Bench is the benchmark for the paper "From Noisy STEM to Crystal Structure: Evidence-Structure CoDiffusion under Composition Constraints" (Chen & You, KDD 2026, Oral), which introduces STEM2Crystal CoDiffusion (SCCD). It evaluates methods that reconstruct a crystal structure from a noisy STEM image when the composition is known. The release has a large synthetic set with controlled noise and a small set of real STEM images, with ground-truth CIFs… See the full description on the dataset page: https://huggingface.co/datasets/gary23ai/STEM2Crystal-Bench.ultrafine-stem-part-2stem-scientific-code-sample
AxiomSet Labs STEM Scientific-Code Sample
A 30-task sample of STEM reasoning and scientific-code problems across five domains.
Domains
Biology: 6 tasks
Chemistry: 6 tasks
Materials Science: 6 tasks
Mathematics: 6 tasks
Physics: 6 tasks
Each task contains two subproblems and one main problem, including prompts, scientific background, testing templates, and reference solutions.
Files
data/sample.jsonl — one task per line; used by the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/AxiomSetLabs/stem-scientific-code-sample.Logics-STEM-SFT-Dataset-Open-5.3Multrafine-stem-part-1gpt-oss-120b-reasoning-STEM-5K
GPT-OSS-120B-Distilled-Reasoning-STEM Dataset
1) Dataset Overview
Data Source Model: gpt-oss-120b-high
Task Type: STEM Reasoning and Problem Solving (Science, Technology, Engineering & Mathematics)
Data Format: `JSON Lines
Fields: generator, category, input, CoT_Native——reasoning, answer
(Consistent with the math dataset, splitting the original 'output' into 'reasoning' and 'answer' for COT/SFT scenarios.)
2) Design Goals (Motivation)
This dataset targets… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120b-reasoning-STEM-5K.salabs-stem-deep-reasoning-cot-v13
🧪 SALabs Multi-Domain STEM Deep Reasoning & Chain-of-Thought (CoT) Corpus (v13.0)
[!IMPORTANT]
💳 Click Here to Purchase Enterprise Commercial License ($2,500 USD) & Instant 31.7MB Master Archive DownloadInstant download of the full lossless master package containing all 1,816 JSONL reasoning records + 13 complete uncompressed text corpora (31.72 MB uncompressed total) + commercial license certificate.
🌟 Executive Summary
The SALabs STEM Deep Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/suitai/salabs-stem-deep-reasoning-cot-v13.Electrical-engineering
To the electrical engineering community
This dataset contains Q&A prompts about electrical engineering, Kicad's EDA software features and scripting console Python codes.
Authors
STEM.AI: stem.ai.mtl@gmail.comWilliam Harbec
stem_zh_instruction
stem_zh_instruction
内容:STEM相关指令(gpt-3.5爬取),包含物理、化学、医学、生物学、地球科学;共计256K条。
Content: STEM related instructions (gpt-3.5 crawled), including physics, chemistry, medicine, biology, and earch science. 256K instruction data in total.
学科 / Subject
文件名 / File Name
数量 / Num
物理 / Physics
phy_50380.json
50,380
化学 / Chemistry
chem_50839.json
50,839
医学 / Medicine
med_54617.json
54,617
生物学 / Biology
bio_50282.json
50,282
地球科学 / Earth Science
earth_50068.json
50,068
总计
256,186… See the full description on the dataset page: https://huggingface.co/datasets/hfl/stem_zh_instruction.MMMLU-STEM-Ko
Details
This is a subset of [openai/MMMLU].
Only the subjects related to STEM were extracted from Korean subset of MMMLU.
The included subjects are
'abstract_algebra',
'anatomy',
'astronomy',
'college_biology',
'college_chemistry',
'college_computer_science',
'college_mathematics',
'college_physics',
'computer_security',
'conceptual_physics',
'electrical_engineering',
'elementary_mathematics',
'high_school_biology',
'high_school_chemistry'… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/MMMLU-STEM-Ko.stem-fixtextbooks_lectures_stemwikiCC-BY-STEMM-Podcast-TranscriptsJosephgflowers__Tinyllama-STEM-Cinder-Agent-v1-details
Dataset Card for Evaluation run of Josephgflowers/Tinyllama-STEM-Cinder-Agent-v1
Dataset automatically created during the evaluation run of model Josephgflowers/Tinyllama-STEM-Cinder-Agent-v1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Josephgflowers__Tinyllama-STEM-Cinder-Agent-v1-details.palladium-stem-preview-25k
⚛️ Palladium-STEM (Preview): High-Density Scientific Corpus
"The Top 0.17% of the Open Web."
Overview
This dataset is a 25,000-document preview of the upcoming Palladium-V2 STEM Corpus. It represents the "Platinum Tier" survivors from a pool of 14.8 million scanned documents, selected for high information density, academic rigor, and reasoning capability.
The "Goldilocks" Methodology
Unlike standard web scrapes, this data was processed using a custom… See the full description on the dataset page: https://huggingface.co/datasets/PalladiumData/palladium-stem-preview-25k.swti-stem-20kswahili-text-corpus
Dataset for Swahili Text Corpus for TTS training
Overview
This dataset contains a synthetic Swahili text corpus designed for training Text-to-Speech (TTS) models. The dataset includes a variety of Swahili phonemes to ensure phonetic diversity and high-quality TTS training.
Statistics
Format: JSONL (JSON Lines)
Data Creation
The dataset was generated using OpenAI's gpt-3.5-turbo model. The model was prompted to produce Swahili sentences that are… See the full description on the dataset page: https://huggingface.co/datasets/stem-content-ai-project/swahili-text-corpus.1.5-Million-English-STEM-Test-Questions-Data-Sample
Description
This dataset contains 1.5 million English science and engineering test questions, including mathematics, physics, chemistry, biology, and other STEM subjects at the university level. Each questions contain title, answer, parse, type, subject, grade. The dataset can be used for large model subject knowledge enhancement tasks.
For more details, please refer to the link: https://www.nexdata.ai/datasets/llm/1881?source=Huggingface
Content
Science subjects… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/1.5-Million-English-STEM-Test-Questions-Data-Sample.stems-predict-dataSTEMmix
STEMmix
Generated with LMDataTools using DataMix.
Samples and combines datasets from Hugging Face.
Here's a thinking process:
Analyze User Input:
Dataset Name: STEMmix
Generated by: DataMix
Sample Entries: Two examples showing conversations between "human" and "gpt". Topics include traffic flow dynamics (chaotic dynamics, car-following models, Optimal Velocity Model) and formal logic/predicate calculus (existential/universal quantifiers, biconditional introduction).… See the full description on the dataset page: https://huggingface.co/datasets/theprint/STEMmix.CC-BY-STEMM-Podcast-Transcripts-2048STEM-AI-mtl_Electrical-engineering-vieMMLU_STEMChosen tasks:
"college_biology",
"college_chemistry",
"college_physics",
"college_computer_science",
"college_mathematics",
"computer_security",
"conceptual_physics",
"electrical_engineering",
"elementary_mathematics",
"high_school_biology",
"high_school_chemistry",
"high_school_computer_science",
"high_school_mathematics",
"high_school_physics",
"high_school_statistics",
"machine_learning",
"astronomy",
"anatomy",
"conceptual_physics",
"abstract_algebra"
STEMScoredTopics-v1.0
Dataset Summary
A synthetic dataset of 5,584 topics, each rated on a 1-5 scale for its relevance to Science, Technology, Engineering, and Mathematics (STEM).
Data Fields
topic: A string representing a topic of study or research.
stemScore: A string from "1" (least STEM) to "5" (most STEM).
Potential Uses
This dataset is useful for a variety of NLP tasks:
Classification: Train a model to classify how STEM-related a given text is.
Instruction… See the full description on the dataset page: https://huggingface.co/datasets/MultivexAI/STEMScoredTopics-v1.0.stemmingstemming-instructionsMuta-STEM-100
Muta STEM 100
Private evaluation dataset containing the fixed 100-prompt Muta STEM battery:
50 mathematics prompts (M01–M50)
50 science prompts (S01–S50)
50 multiple-choice and 50 structured-written prompts
Each row contains id, title, subject, format, text, expected, source, and suite.
This is an evaluation artifact, not a training split. Publishing or training on it would contaminate future benchmark results. Model responses are not included.
Source artifact SHA256:… See the full description on the dataset page: https://huggingface.co/datasets/timiiowolabi/Muta-STEM-100.pgtr_synth_doc_stem
