CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lhpku20010120 /K12-KGraph K12-KGraph This repository contains the dataset release for the paper "K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs". Paper | Project page | Code Overview K12-KGraph is a curriculum-aligned knowledge graph built from official People's Education Press (PEP) K-12 textbooks. It focuses on curriculum cognition, namely the structured understanding of how school knowledge is organized, connected, and sequenced. The… See the full description on the dataset page: https://huggingface.co/datasets/lhpku20010120/K12-KGraph.text-generation16 likes2.6k downloads2mo agoHugging Face02lfaviate /China-K12-STEM-10K-CoT-Reasoning K12-STEM-CoT-Chinese 1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams. The largest structured Chinese math/physics/chemistry reasoning dataset. This is a curated sample (10,000 problems) of the full 1.54M dataset available via API. Full Dataset Access Access the full 1,540,000+ problems via API → This Sample Full API Total problems 10,025 1,540,000+ With CoT solutions 10,025 1,490,000+ With diagrams 6,093 740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.tabularquestion-answering10K<n<100K3 likes604 downloads7mo agoHugging Face03marin-community /openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16 OpenThoughts-4 Code SDG: Qwen3-30B-A3B-Thinking-2507 (n=16, top-16 logprobs) Synthetic generations from Qwen/Qwen3-30B-A3B-Thinking-2507 on the Marin OpenThoughts-4 code SDG prompt set. Each prompt is sampled n=16 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes465 downloads5mo agoHugging Face04marin-community /openthoughts4-code-9168-prompts-qwen3-32b-n16-flattened-logprobs-k16 OpenThoughts-4 Code SDG: Qwen3-32B (n=16, top-16 logprobs) Synthetic generations from Qwen/Qwen3-32B on the Marin OpenThoughts-4 code SDG prompt set. Each prompt is sampled n=16 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-32b-n16-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes445 downloads5mo agoHugging Face05marin-community /openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-30B-A3B-Thinking-2507 (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-30B-A3B-Thinking-2507 on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes227 downloads5mo agoHugging Face06marin-community /openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-32B (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-32B on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes212 downloads5mo agoHugging Face07marin-community /openthoughts4-code-9168-prompts-qwen3-4b-n16-flattened-logprobs-k16 OpenThoughts-4 Code SDG: Qwen3-4B (n=16, top-16 logprobs) Synthetic generations from Qwen/Qwen3-4B on the Marin OpenThoughts-4 code SDG prompt set. Each prompt is sampled n=16 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model Qwen/Qwen3-4B… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-4b-n16-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes161 downloads5mo agoHugging Face08robworks-software /k12-standards-instruction-tasks K-12 Curriculum Tasks (generated) 2,489 generated instruction/input/output records covering five curriculum tasks: assessment creation, learning objective generation, misconception detection, standard explanation, and standards Q&A. Content is predominantly mathematics. Important: the name is misleading Despite the name, this dataset contains no school directory data. There are four columns - task, input, output, metadata - and no staff, principal, or school… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-standards-instruction-tasks.texttext-generation1K<n<10K1 likes150 downloads2mo agoHugging Face09FoundryAILabs /k12-indian-curriculum-4.9m BharatLLM K-12 Indian Curriculum Dataset (4.9M) 4,904,936 question-answer pairs covering CBSE/NCERT K-12 curriculum across 12 Indian languages. Language Script Entries English Latin ~594K Hindi Devanagari ~449K Bengali Bengali ~408K Telugu Telugu ~408K Tamil Tamil ~408K Kannada Kannada ~408K Malayalam Malayalam ~408K Marathi Devanagari ~408K Gujarati Gujarati ~408K Odia Odia ~408K PunjabiGurmukhi ~408K Urdu Nastaliq ~374K Format {… See the full description on the dataset page: https://huggingface.co/datasets/FoundryAILabs/k12-indian-curriculum-4.9m.textquestion-answering1M<n<10M1 likes114 downloads6mo agoHugging Face10krittus /k12-indian-curriculum-4.9m BharatLLM K-12 Indian Curriculum Dataset (4.9M) 4,904,936 question-answer pairs covering CBSE/NCERT K-12 curriculum across 12 Indian languages. Language Script Entries English Latin ~594K Hindi Devanagari ~449K Bengali Bengali ~408K Telugu Telugu ~408K Tamil Tamil ~408K Kannada Kannada ~408K Malayalam Malayalam ~408K Marathi Devanagari ~408K Gujarati Gujarati ~408K Odia Odia ~408K Punjabi Gurmukhi ~408K Urdu Nastaliq ~374K Format… See the full description on the dataset page: https://huggingface.co/datasets/krittus/k12-indian-curriculum-4.9m.textquestion-answering1M<n<10M1 likes94 downloads6mo agoHugging Face11opendatalab /K12textbook覆盖小学、初中、高中的高质量中文K12教材语料,经过精细的文本抽取和数据处理,可用于学术研究 texttext-generationn<1K3 likes88 downloads1y agoHugging Face12SchoolData /us-k12-schools US K–12 Schools — Open Dataset for the AI Era One encyclopedic paragraph for every one of the 122,675 K–12 schools in the United States, ready to use as pretraining text. Ask a language model about a large suburban high school and it will answer. Ask it about the K–8 school in a rural county of 4,000 people and it has nothing to say — because nothing about that school was ever written down on the open web. Only about 12% of American schools have a Wikipedia article at all, and… See the full description on the dataset page: https://huggingface.co/datasets/SchoolData/us-k12-schools.texttext-generation100K<n<1M0 likes84 downloads18d agoHugging Face13K1zE /BPM BPM Training Prompts (mix-20k) arXiv:2607.22334 · Project page Prompt corpus behind every reported result in BPM (Byte-Prefix Marginalization), a cross-tokenizer on-policy distillation method. Prompt-only. Domain Rows Source Upstream license Mathematics 10,000 dapo-17k Apache-2.0 Code 10,000 taco:* Apache-2.0 Fields Field Type Notes prompt list of {role, content} Chat-format user turn label string Reference answer for mathematics;… See the full description on the dataset page: https://huggingface.co/datasets/K1zE/BPM.texttext-generation10K<n<100K0 likes67 downloads2mo agoHugging Face14zl2023 /K12-KGraph K12-KGraph This repository contains the dataset release for the paper "K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs". Paper | Project page | Code Overview K12-KGraph is a curriculum-aligned knowledge graph built from official People's Education Press (PEP) K-12 textbooks. It focuses on curriculum cognition, namely the structured understanding of how school knowledge is organized, connected, and sequenced. The… See the full description on the dataset page: https://huggingface.co/datasets/zl2023/K12-KGraph.text-generation0 likes61 downloads2mo agoHugging Face15a13905873166 /China-K12-STEM-10K-CoT-Reasoning K12-STEM-CoT-Chinese 1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams. The largest structured Chinese math/physics/chemistry reasoning dataset. This is a curated sample (10,000 problems) of the full 1.54M dataset available via API. Full Dataset Access Access the full 1,540,000+ problems via API → This Sample Full API Total problems 10,025 1,540,000+ With CoT solutions 10,025 1,490,000+ With diagrams 6,093 740… See the full description on the dataset page: https://huggingface.co/datasets/a13905873166/China-K12-STEM-10K-CoT-Reasoning.tabularquestion-answering10K<n<100K1 likes58 downloads10d agoHugging Face16marin-community /openthoughts4-science-26041-prompts-qwen3-4b-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-4B (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-4B on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-4b-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes57 downloads5mo agoHugging Face17robworks-software /k12-mathematics-standards-expanded K-12 Mathematics Standards, expanded (generated instruction data) 4,965 instruction/input/output records for mathematics, generated around a K-12 standards taxonomy for instruction-tuning and educational-content experiments. How this was built (read this first) These are programmatically generated training examples, not curriculum written by educators and not the text of any official standard. A generator combined standards metadata - codes, grade levels, domains… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-mathematics-standards-expanded.texttext-generation1K<n<10K0 likes50 downloads2mo agoHugging Face18robworks-software /k12-ela-standards-expanded K-12 ELA Standards, expanded (generated instruction data) 12,282 instruction/input/output records for English Language Arts, generated around a K-12 standards taxonomy for instruction-tuning and educational-content experiments. How this was built (read this first) These are programmatically generated training examples, not curriculum written by educators and not the text of any official standard. A generator combined standards metadata - codes, grade levels… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-ela-standards-expanded.texttext-generation10K<n<100K0 likes50 downloads2mo agoHugging Face19robworks-software /k12-science-standards [!WARNING] Deprecated - use k12-science-standards-expanded instead. This dataset is superseded: every instruction in this set also appears there, plus 1,123 more and nine additional metadata columns. Nothing here is unique to it. It stays online so existing references keep resolving, but it will not be updated. New work should point at robworks-software/k12-science-standards-expanded. K-12 Science Standards (generated instruction data) 6,787 instruction/input/output records… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-science-standards.texttext-classification1K<n<10K0 likes49 downloads2mo agoHugging Face20robworks-software /texas-k12-curriculum-standards-teks Texas K-12 Curriculum Standards (TEKS-derived) 15,040 generated learning-objective records organized around the Texas Essential Knowledge and Skills (TEKS) taxonomy, spanning core academic subjects, Career & Technical Education clusters, and specialized program areas. How this was built (read this first) These records are programmatically generated, not transcribed from official standards documents. A generator took a standards taxonomy - codes, grade levels… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/texas-k12-curriculum-standards-teks.texttext-classification10K<n<100K0 likes44 downloads2mo agoHugging Face21robworks-software /k12-mathematics-standards-aligned [!WARNING] Deprecated - use k12-mathematics-standards-expanded instead. This dataset is superseded: every input in this set also appears there, plus 366 more and two additional metadata columns. Nothing here is unique to it. It stays online so existing references keep resolving, but it will not be updated. New work should point at robworks-software/k12-mathematics-standards-expanded. K-12 Mathematics Standards (generated instruction data) 4,397 instruction/input/output records… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-mathematics-standards-aligned.texttext-generation1K<n<10K0 likes42 downloads2mo agoHugging Face22robworks-software /k12-social-studies-standards K-12 Social Studies Standards (generated instruction data) 15,982 instruction/input/output records for social studies (civics, history, geography, economics), generated around a K-12 standards taxonomy for instruction-tuning and educational-content experiments. How this was built (read this first) These are programmatically generated training examples, not curriculum written by educators and not the text of any official standard. A generator combined standards… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-social-studies-standards.textquestion-answering10K<n<100K0 likes41 downloads2mo agoHugging Face23robworks-software /k12-business-economics-standards K-12 Business and Economics Standards 1,236 generated learning-objective records covering financial literacy, personal finance, entrepreneurship, business management, and career development, organized around the Jump$tart Personal Financial Education and NBEA Business Education standard structures. How this was built (read this first) These records are programmatically generated, not transcribed from official standards documents. A generator took a standards… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-business-economics-standards.texttext-classification1K<n<10K0 likes38 downloads2mo agoHugging Face24robworks-software /k12-ela-standards [!WARNING] Deprecated - use k12-ela-standards-expanded instead. This dataset is superseded: every input in this set also appears there, plus 1,433 more and five additional metadata columns. Nothing here is unique to it. It stays online so existing references keep resolving, but it will not be updated. New work should point at robworks-software/k12-ela-standards-expanded. K-12 ELA Standards (generated instruction data) 6,487 instruction/input/output records for English Language… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-ela-standards.texttext-generation1K<n<10K0 likes34 downloads2mo agoHugging Face25ijanhq8809 /k12-ela-standards K-12 English Language Arts Standards Dataset 📚 Comprehensive ELA Education Dataset This dataset provides complete K-12 English Language Arts standards coverage for AI training and educational technology development. 📊 Dataset Summary The K-12 ELA Standards Dataset contains: 546 educational standards across K-12 6487 AI training samples Complete coverage of all major ELA domains 🎓 Educational Coverage 📖 Reading Literature: Fiction, poetry… See the full description on the dataset page: https://huggingface.co/datasets/ijanhq8809/k12-ela-standards.texttext-generation1K<n<10K1 likes32 downloads9mo agoHugging Face26lemonteaa /kindergarten_K1_english_lesson_autogentexttext-generationn<1K0 likes24 downloads2y agoHugging Face27robworks-software /k12-science-standards-expanded K-12 Science Standards, expanded (generated instruction data) 15,354 instruction/input/output records for science, generated around a K-12 standards taxonomy for instruction-tuning and educational-content experiments. How this was built (read this first) These are programmatically generated training examples, not curriculum written by educators and not the text of any official standard. A generator combined standards metadata - codes, grade levels, domains… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-science-standards-expanded.texttext-classification10K<n<100K0 likes23 downloads2mo agoHugging Face28aikalvi /edu-k12-instruct K-12 Educational Instruction Dataset A curated dataset for fine-tuning language models on K-12 educational content. Dataset Description This dataset contains question-answer pairs covering educational topics for grades 3-12, designed for instruction-tuning language models to be helpful educational assistants. Subjects Covered Mathematics (284 examples) - Arithmetic, algebra, geometry, basic math concepts English (238 examples) - Grammar, vocabulary, writing… See the full description on the dataset page: https://huggingface.co/datasets/aikalvi/edu-k12-instruct.textquestion-answeringn<1K0 likes23 downloads8mo agoHugging Face29CL-From-Nothing /code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288 code_rose_initial_1_7B_SFT_10K — rollouts (Qwen3-4B-Thinking-2507, k=12) Pass@k completions generated with vLLM over the prefixes in CL-From-Nothing/code_rose_initial_1_7B_SFT_10K. Generation config Model Qwen3-4B-Thinking-2507 Samples per question (k) 12 Temperature 0.7 top_p 0.9 max_tokens 12288 max_model_len 32768 Questions 7250 (index 0–7249, full split) Total rows 87000 (7250 × 12) Generated by complete_prefix_vllm.py… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288.tabulartext-generation10K<n<100K0 likes22 downloads3mo agoHugging Face30tttonyyy /NMC-cn_k12-20k-r1_32b_distilled本数据集数据来源为NuminaMath-CoT数据集的cn_k12数据。我们从这里面提取了20000条问题,并使用DeepSeek-R1-Distill-Qwen-32B模型进行了回答。 distilled_s0_e20000.jsonl包含这个数据集的数据,下面介绍数据标签: idx:索引号(0~19999) question:原数据集中的problem标签,是一个可能包含多个子问题的数学问题字符串 gt_cot:愿数据集中的solution标签,是经过GPT-4o整理的答案字符串 pred_cot:根据question标签,模型DeepSeek-R1-Distill-Qwen-32B的回答字符串 pred_cot_token_len:pred_cot标签下的字符串转化成token之后的长度(不包含最前面的<think>\n部分,这个在生成的时候是在prompt里面,我后来加到这里了) message:根据question标签和pred_cot标签,构造的问题-回答数据对 统计了一下平均回答token长度,为3169.4251 tabulartext-generation10K<n<100K0 likes20 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.