datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
China-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.k12-standards-instruction-tasks
K-12 Curriculum Tasks (generated)
2,489 generated instruction/input/output records covering five curriculum tasks:
assessment creation, learning objective generation, misconception detection, standard
explanation, and standards Q&A. Content is predominantly mathematics.
Important: the name is misleading
Despite the name, this dataset contains no school directory data. There are four
columns - task, input, output, metadata - and no staff, principal, or school… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-standards-instruction-tasks.k12-indian-curriculum-4.9m
BharatLLM K-12 Indian Curriculum Dataset (4.9M)
4,904,936 question-answer pairs covering CBSE/NCERT K-12 curriculum across 12 Indian languages.
Language
Script
Entries
English
Latin
~594K
Hindi
Devanagari
~449K
Bengali
Bengali
~408K
Telugu
Telugu
~408K
Tamil
Tamil
~408K
Kannada
Kannada
~408K
Malayalam
Malayalam
~408K
Marathi
Devanagari
~408K
Gujarati
Gujarati
~408K
Odia
Odia
~408K
PunjabiGurmukhi
~408K
Urdu
Nastaliq
~374K
Format
{… See the full description on the dataset page: https://huggingface.co/datasets/FoundryAILabs/k12-indian-curriculum-4.9m.k12-indian-curriculum-4.9m
BharatLLM K-12 Indian Curriculum Dataset (4.9M)
4,904,936 question-answer pairs covering CBSE/NCERT K-12 curriculum across 12 Indian languages.
Language
Script
Entries
English
Latin
~594K
Hindi
Devanagari
~449K
Bengali
Bengali
~408K
Telugu
Telugu
~408K
Tamil
Tamil
~408K
Kannada
Kannada
~408K
Malayalam
Malayalam
~408K
Marathi
Devanagari
~408K
Gujarati
Gujarati
~408K
Odia
Odia
~408K
Punjabi
Gurmukhi
~408K
Urdu
Nastaliq
~374K
Format… See the full description on the dataset page: https://huggingface.co/datasets/krittus/k12-indian-curriculum-4.9m.K12textbook覆盖小学、初中、高中的高质量中文K12教材语料,经过精细的文本抽取和数据处理,可用于学术研究
us-k12-schools
US K–12 Schools — Open Dataset for the AI Era
One encyclopedic paragraph for every one of the 122,675 K–12 schools in the United
States, ready to use as pretraining text.
Ask a language model about a large suburban high school and it will answer. Ask it about
the K–8 school in a rural county of 4,000 people and it has nothing to say — because
nothing about that school was ever written down on the open web. Only about 12% of
American schools have a Wikipedia article at all, and… See the full description on the dataset page: https://huggingface.co/datasets/SchoolData/us-k12-schools.China-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740… See the full description on the dataset page: https://huggingface.co/datasets/a13905873166/China-K12-STEM-10K-CoT-Reasoning.k12-mathematics-standards-expanded
K-12 Mathematics Standards, expanded (generated instruction data)
4,965 instruction/input/output records for mathematics, generated around a K-12
standards taxonomy for instruction-tuning and educational-content experiments.
How this was built (read this first)
These are programmatically generated training examples, not curriculum written by
educators and not the text of any official standard. A generator combined standards
metadata - codes, grade levels, domains… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-mathematics-standards-expanded.k12-science-standards
[!WARNING]
Deprecated - use k12-science-standards-expanded instead.
This dataset is superseded: every instruction in this set also appears there, plus 1,123 more and nine additional metadata columns. Nothing here is unique to it.
It stays online so existing references keep resolving, but it will not be updated.
New work should point at robworks-software/k12-science-standards-expanded.
K-12 Science Standards (generated instruction data)
6,787 instruction/input/output records… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-science-standards.k12-ela-standards-expanded
K-12 ELA Standards, expanded (generated instruction data)
12,282 instruction/input/output records for English Language Arts, generated around a K-12
standards taxonomy for instruction-tuning and educational-content experiments.
How this was built (read this first)
These are programmatically generated training examples, not curriculum written by
educators and not the text of any official standard. A generator combined standards
metadata - codes, grade levels… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-ela-standards-expanded.k12-mathematics-standards-aligned
[!WARNING]
Deprecated - use k12-mathematics-standards-expanded instead.
This dataset is superseded: every input in this set also appears there, plus 366 more and two additional metadata columns. Nothing here is unique to it.
It stays online so existing references keep resolving, but it will not be updated.
New work should point at robworks-software/k12-mathematics-standards-expanded.
K-12 Mathematics Standards (generated instruction data)
4,397 instruction/input/output records… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-mathematics-standards-aligned.texas-k12-curriculum-standards-teks
Texas K-12 Curriculum Standards (TEKS-derived)
15,040 generated learning-objective records organized around the Texas Essential
Knowledge and Skills (TEKS) taxonomy, spanning core academic subjects, Career & Technical
Education clusters, and specialized program areas.
How this was built (read this first)
These records are programmatically generated, not transcribed from official standards
documents. A generator took a standards taxonomy - codes, grade levels… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/texas-k12-curriculum-standards-teks.k12-social-studies-standards
K-12 Social Studies Standards (generated instruction data)
15,982 instruction/input/output records for social studies (civics, history, geography, economics), generated around a K-12
standards taxonomy for instruction-tuning and educational-content experiments.
How this was built (read this first)
These are programmatically generated training examples, not curriculum written by
educators and not the text of any official standard. A generator combined standards… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-social-studies-standards.k12-business-economics-standards
K-12 Business and Economics Standards
1,236 generated learning-objective records covering financial literacy, personal finance,
entrepreneurship, business management, and career development, organized around the
Jump$tart Personal Financial Education and NBEA Business Education standard structures.
How this was built (read this first)
These records are programmatically generated, not transcribed from official standards
documents. A generator took a standards… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-business-economics-standards.k12-ela-standards
[!WARNING]
Deprecated - use k12-ela-standards-expanded instead.
This dataset is superseded: every input in this set also appears there, plus 1,433 more and five additional metadata columns. Nothing here is unique to it.
It stays online so existing references keep resolving, but it will not be updated.
New work should point at robworks-software/k12-ela-standards-expanded.
K-12 ELA Standards (generated instruction data)
6,487 instruction/input/output records for English Language… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-ela-standards.k12-ela-standards
K-12 English Language Arts Standards Dataset
📚 Comprehensive ELA Education Dataset
This dataset provides complete K-12 English Language Arts standards coverage for AI training and educational technology development.
📊 Dataset Summary
The K-12 ELA Standards Dataset contains:
546 educational standards across K-12
6487 AI training samples
Complete coverage of all major ELA domains
🎓 Educational Coverage
📖 Reading Literature: Fiction, poetry… See the full description on the dataset page: https://huggingface.co/datasets/ijanhq8809/k12-ela-standards.edu-k12-instruct
K-12 Educational Instruction Dataset
A curated dataset for fine-tuning language models on K-12 educational content.
Dataset Description
This dataset contains question-answer pairs covering educational topics for grades 3-12, designed for instruction-tuning language models to be helpful educational assistants.
Subjects Covered
Mathematics (284 examples) - Arithmetic, algebra, geometry, basic math concepts
English (238 examples) - Grammar, vocabulary, writing… See the full description on the dataset page: https://huggingface.co/datasets/aikalvi/edu-k12-instruct.code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288
code_rose_initial_1_7B_SFT_10K — rollouts (Qwen3-4B-Thinking-2507, k=12)
Pass@k completions generated with vLLM over the prefixes in
CL-From-Nothing/code_rose_initial_1_7B_SFT_10K.
Generation config
Model
Qwen3-4B-Thinking-2507
Samples per question (k)
12
Temperature
0.7
top_p
0.9
max_tokens
12288
max_model_len
32768
Questions
7250 (index 0–7249, full split)
Total rows
87000 (7250 × 12)
Generated by complete_prefix_vllm.py… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288.k12-science-standards-expanded
K-12 Science Standards, expanded (generated instruction data)
15,354 instruction/input/output records for science, generated around a K-12
standards taxonomy for instruction-tuning and educational-content experiments.
How this was built (read this first)
These are programmatically generated training examples, not curriculum written by
educators and not the text of any official standard. A generator combined standards
metadata - codes, grade levels, domains… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-science-standards-expanded.california-k12-standards
California K-12 Educational Standards
3,410 records organized around California K-12 standards frameworks, including Common
Core, NGSS, ELD, CTE, and Ethnic Studies. Records carry a standard identifier, grade
level, subject area, domain, and generated learning-objective and application text.
How this was built (read this first)
These records are programmatically generated, not transcribed from official standards
documents. A generator took a standards taxonomy -… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/california-k12-standards.NMC-cn_k12-20k-r1_32b_distilled本数据集数据来源为NuminaMath-CoT数据集的cn_k12数据。我们从这里面提取了20000条问题,并使用DeepSeek-R1-Distill-Qwen-32B模型进行了回答。
distilled_s0_e20000.jsonl包含这个数据集的数据,下面介绍数据标签:
idx:索引号(0~19999)
question:原数据集中的problem标签,是一个可能包含多个子问题的数学问题字符串
gt_cot:愿数据集中的solution标签,是经过GPT-4o整理的答案字符串
pred_cot:根据question标签,模型DeepSeek-R1-Distill-Qwen-32B的回答字符串
pred_cot_token_len:pred_cot标签下的字符串转化成token之后的长度(不包含最前面的<think>\n部分,这个在生成的时候是在prompt里面,我后来加到这里了)
message:根据question标签和pred_cot标签,构造的问题-回答数据对
统计了一下平均回答token长度,为3169.4251
K12textbook覆盖小学、初中、高中的高质量中文K12教材语料,经过精细的文本抽取和数据处理,可用于学术研究
