CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lhpku20010120 /K12-KGraph K12-KGraph This repository contains the dataset release for the paper "K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs". Paper | Project page | Code Overview K12-KGraph is a curriculum-aligned knowledge graph built from official People's Education Press (PEP) K-12 textbooks. It focuses on curriculum cognition, namely the structured understanding of how school knowledge is organized, connected, and sequenced. The… See the full description on the dataset page: https://huggingface.co/datasets/lhpku20010120/K12-KGraph.text-generation16 likes2.5k downloads2mo agoHugging Face02aslawliet /cn-k12text100K<n<1M13 likes2.4k downloads2y agoHugging Face03alexkstern /dyck-k128-seq_len_2048-1B dyck-k128-seq_len_2048-1B Procedurally generated k-shuffle Dyck bracket sequences (Hu et al. 2025, arXiv:2502.19249), as flat uint16 token-id .bin files. Token ids are 0-based: opening bracket type i is id i and its matching close is i + k, so ids span [0, 2k) and the vocabulary is 2k = 256. Grammar parameters param value k (bracket types) 128 max_depth 16 p_open 0.5 seq_length 2048 file split tokens train.bin train 999,999,488 val.bin val 10,000… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/dyck-k128-seq_len_2048-1B.tabularn<1K0 likes1.5k downloads4mo agoHugging Face04tunaaa126 /K12-Dataset K12-KGraph K12-KGraph is a curriculum-aligned knowledge graph built from official People's Education Press (PEP) K-12 textbooks. It focuses on curriculum cognition, namely the structured understanding of how school knowledge is organized, connected, and sequenced. The current release covers mathematics, physics, chemistry, and biology across primary, middle, and high school, and includes three resources derived from the same graph: K12-KGraph: the core knowledge graph K12-Bench: a… See the full description on the dataset page: https://huggingface.co/datasets/tunaaa126/K12-Dataset.text1K<n<10K0 likes682 downloads5mo agoHugging Face05lfaviate /China-K12-STEM-10K-CoT-Reasoning K12-STEM-CoT-Chinese 1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams. The largest structured Chinese math/physics/chemistry reasoning dataset. This is a curated sample (10,000 problems) of the full 1.54M dataset available via API. Full Dataset Access Access the full 1,540,000+ problems via API → This Sample Full API Total problems 10,025 1,540,000+ With CoT solutions 10,025 1,490,000+ With diagrams 6,093 740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.tabularquestion-answering10K<n<100K3 likes605 downloads7mo agoHugging Face06zhenliuu /k12-multidisciplinary K12 多学科图文推理数据集 面向中小学数学、物理、生物、地理和化学的多学科图文推理数据。 GitHub 主页与训练代码 数据范围 配置 split 题目数 含图题数 唯一图片数 default raw 735,650 515,089 416,099 dapo train 1,160 712 719 原始数据与训练集按用途分别提供。训练集从总数据集中筛选整理而来,面向数学与物理推理任务,可用于不同模型与训练框架。 原始数据 数据由公开开源数据筛选、自建实体书OCR抽取及基于vLLM的合成与改写三部分构成,经过图片回收与校验、结构统一、来源标签清理、重复题处理及答案冲突复核。数据包含纯文本题和含图题,同时覆盖选择题与非选择题。 数学62,966题、物理164,398题、生物199,084题、地理174,282题、化学134,920题。含图题占70.02%。 数学与物理训练集… See the full description on the dataset page: https://huggingface.co/datasets/zhenliuu/k12-multidisciplinary.documentvisual-question-answering100K<n<1M0 likes219 downloads3d agoHugging Face07Neelectric /OpenR1-Math-cn_k12-91ktext10K<n<100K0 likes201 downloads1y agoHugging Face08xyliu6 /k12-freeformimage10K<n<100K2 likes187 downloads1y agoHugging Face09robworks-software /us-k12-schools-directory US K-12 Schools Directory A directory of 124,613 US K-12 schools covering all 50 states, DC, and US territories, compiled from federal and state government sources. Each record carries directory information (address, phone, website), enrollment and demographics, and, where a source supplied it, a principal name and email. This is a compilation of public government data. It is not a survey, and no field was independently verified against the school itself. Loading… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/us-k12-schools-directory.tabulartabular-classification100K<n<1M0 likes165 downloads2mo agoHugging Face10robworks-software /k12-standards-instruction-tasks K-12 Curriculum Tasks (generated) 2,489 generated instruction/input/output records covering five curriculum tasks: assessment creation, learning objective generation, misconception detection, standard explanation, and standards Q&A. Content is predominantly mathematics. Important: the name is misleading Despite the name, this dataset contains no school directory data. There are four columns - task, input, output, metadata - and no staff, principal, or school… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-standards-instruction-tasks.texttext-generation1K<n<10K1 likes147 downloads2mo agoHugging Face11anonymous-K12 /K12-KGraph K12-KGraph This repository contains the dataset release for the paper "K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs". Overview K12-KGraph is a curriculum-aligned knowledge graph built from official People's Education Press (PEP) K-12 textbooks. It focuses on curriculum cognition, namely the structured understanding of how school knowledge is organized, connected, and sequenced. The current release covers mathematics… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-K12/K12-KGraph.0 likes144 downloads5mo agoHugging Face12FoundryAILabs /k12-indian-curriculum-4.9m BharatLLM K-12 Indian Curriculum Dataset (4.9M) 4,904,936 question-answer pairs covering CBSE/NCERT K-12 curriculum across 12 Indian languages. Language Script Entries English Latin ~594K Hindi Devanagari ~449K Bengali Bengali ~408K Telugu Telugu ~408K Tamil Tamil ~408K Kannada Kannada ~408K Malayalam Malayalam ~408K Marathi Devanagari ~408K Gujarati Gujarati ~408K Odia Odia ~408K PunjabiGurmukhi ~408K Urdu Nastaliq ~374K Format {… See the full description on the dataset page: https://huggingface.co/datasets/FoundryAILabs/k12-indian-curriculum-4.9m.textquestion-answering1M<n<10M1 likes112 downloads6mo agoHugging Face13WeiXiCZ /cast-k12-mlp-stage2-cache-v1 CAST K=12 MLP Stage-II caches Frozen slot/action caches for research experiments. Both sources were encoded with the same NAVSIM-trained, no-geometry Stage-I checkpoint. NAVSIM records: 97,055 WorldEngine in-grid records: 8,635 Shape: history 4 + future 8, K=12, slot dimension=256 Slot encoder SHA-256: 3f6f2b885ff6acab00f3a025c3dfd924040f3e2592e5369484ca779a2ae125dc Action tokenizer hash: 705d7abd9d18122d The many small PyTorch records are stored as independently downloadable… See the full description on the dataset page: https://huggingface.co/datasets/WeiXiCZ/cast-k12-mlp-stage2-cache-v1.other0 likes111 downloads22d agoHugging Face14krittus /k12-indian-curriculum-4.9m BharatLLM K-12 Indian Curriculum Dataset (4.9M) 4,904,936 question-answer pairs covering CBSE/NCERT K-12 curriculum across 12 Indian languages. Language Script Entries English Latin ~594K Hindi Devanagari ~449K Bengali Bengali ~408K Telugu Telugu ~408K Tamil Tamil ~408K Kannada Kannada ~408K Malayalam Malayalam ~408K Marathi Devanagari ~408K Gujarati Gujarati ~408K Odia Odia ~408K Punjabi Gurmukhi ~408K Urdu Nastaliq ~374K Format… See the full description on the dataset page: https://huggingface.co/datasets/krittus/k12-indian-curriculum-4.9m.textquestion-answering1M<n<10M1 likes98 downloads6mo agoHugging Face15WeiXiCZ /cast-k12-mlp-stage2-geometry-cache-v1 CAST K=12 MLP Stage-II caches Frozen slot/action caches for research experiments. Both sources were encoded with the same NAVSIM-trained, geometry Stage-I checkpoint. NAVSIM records: 97,055 WorldEngine in-grid records: 8,635 Shape: history 4 + future 8, K=12, slot dimension=256 Slot encoder SHA-256: 951fe094569cc54ae18dace20d5d8e1d2a82177cac9e97310c1f4b28c0a657d7 Action tokenizer hash: 705d7abd9d18122d The many small PyTorch records are stored as independently downloadable… See the full description on the dataset page: https://huggingface.co/datasets/WeiXiCZ/cast-k12-mlp-stage2-geometry-cache-v1.other0 likes90 downloads20d agoHugging Face16opendatalab /K12textbook覆盖小学、初中、高中的高质量中文K12教材语料,经过精细的文本抽取和数据处理,可用于学术研究 texttext-generationn<1K3 likes85 downloads1y agoHugging Face17SchoolData /us-k12-schools US K–12 Schools — Open Dataset for the AI Era One encyclopedic paragraph for every one of the 122,675 K–12 schools in the United States, ready to use as pretraining text. Ask a language model about a large suburban high school and it will answer. Ask it about the K–8 school in a rural county of 4,000 people and it has nothing to say — because nothing about that school was ever written down on the open web. Only about 12% of American schools have a Wikipedia article at all, and… See the full description on the dataset page: https://huggingface.co/datasets/SchoolData/us-k12-schools.texttext-generation100K<n<1M0 likes80 downloads17d agoHugging Face18Cierra0506 /MM-K12 MM-K12 [📂 GitHub] [📜 Paper] MM-K12 is a curated, high-quality dataset containing 10,000 multimodal math problems sourced from K-12 educational content. Each problem includes both textual and visual components, covering a wide range of mathematical topics (e.g., arithmetic, geometry, algebra). All problems have unique, verifiable answers, making the dataset ideal for supervised training, evaluation, and reward modeling in multimodal mathematical reasoning tasks. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Cierra0506/MM-K12.imageimage-text-to-text10K<n<100K5 likes76 downloads1y agoHugging Face19lipku1999 /K12-Vista3 likes68 downloads1y agoHugging Face20xyliu6 /k12-freeform-extendedimage10K<n<100K1 likes59 downloads1y agoHugging Face21zl2023 /K12-KGraph K12-KGraph This repository contains the dataset release for the paper "K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs". Paper | Project page | Code Overview K12-KGraph is a curriculum-aligned knowledge graph built from official People's Education Press (PEP) K-12 textbooks. It focuses on curriculum cognition, namely the structured understanding of how school knowledge is organized, connected, and sequenced. The… See the full description on the dataset page: https://huggingface.co/datasets/zl2023/K12-KGraph.text-generation0 likes58 downloads1mo agoHugging Face22a13905873166 /China-K12-STEM-10K-CoT-Reasoning K12-STEM-CoT-Chinese 1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams. The largest structured Chinese math/physics/chemistry reasoning dataset. This is a curated sample (10,000 problems) of the full 1.54M dataset available via API. Full Dataset Access Access the full 1,540,000+ problems via API → This Sample Full API Total problems 10,025 1,540,000+ With CoT solutions 10,025 1,490,000+ With diagrams 6,093 740… See the full description on the dataset page: https://huggingface.co/datasets/a13905873166/China-K12-STEM-10K-CoT-Reasoning.tabularquestion-answering10K<n<100K1 likes53 downloads9d agoHugging Face23tttonyyy /NuminaMath-CoT-cn_k12-20000-old本数据集为NuminaMath-CoT在cn_k12类别下前20000条数据中进行的长CoT文本生成尝试。 里面有两个数据文件,completion.jsonl和distilled.jsonl。 其中,completion.jsonl是gemma2-27b-it生成的数据,因为gemma2-27b-it最后得到的结果可能提取不出\boxed{}里面的内容等因素,所以最后生成了约18.6k的数据。 而distilled.jsonl是基于上面的回答成功的问题,使用DeepSeek-R1-Distill-Qwen-32B生成的数据(其实不应该这样,之后有机会再从20000条里面直接生成)。 因为NuminaMath-CoT的cn_k12数据大多是一些综合应用题,一条数据中会有多个子问题,回答的时候也要分别回答。对问题回答的答案进行验证较难,尚未进行答案验证,需要下载下来之后再进行更精细化的处理。 completion.jsonl中的数据标签描述: source (String):数据来源类别,为原始标签 problem (String):数学问题文本,为原始标签 solution… See the full description on the dataset page: https://huggingface.co/datasets/tttonyyy/NuminaMath-CoT-cn_k12-20000-old.question-answering10K<n<100K0 likes52 downloads2y agoHugging Face24robworks-software /k12-mathematics-standards-expanded K-12 Mathematics Standards, expanded (generated instruction data) 4,965 instruction/input/output records for mathematics, generated around a K-12 standards taxonomy for instruction-tuning and educational-content experiments. How this was built (read this first) These are programmatically generated training examples, not curriculum written by educators and not the text of any official standard. A generator combined standards metadata - codes, grade levels, domains… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-mathematics-standards-expanded.texttext-generation1K<n<10K0 likes52 downloads2mo agoHugging Face25robworks-software /k12-science-standards [!WARNING] Deprecated - use k12-science-standards-expanded instead. This dataset is superseded: every instruction in this set also appears there, plus 1,123 more and nine additional metadata columns. Nothing here is unique to it. It stays online so existing references keep resolving, but it will not be updated. New work should point at robworks-software/k12-science-standards-expanded. K-12 Science Standards (generated instruction data) 6,787 instruction/input/output records… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-science-standards.texttext-classification1K<n<10K0 likes52 downloads2mo agoHugging Face26zhangtao00001 /K12Vista0 likes51 downloads1y agoHugging Face27ritwika96 /chess-explained-sf6-k12-curriculum-500ktext100K<n<1M0 likes48 downloads8mo agoHugging Face28Neelectric /OpenR1-Math-220k_CN-K12_OLMo-2_4096tokstext10K<n<100K0 likes46 downloads1y agoHugging Face29robworks-software /k12-ela-standards-expanded K-12 ELA Standards, expanded (generated instruction data) 12,282 instruction/input/output records for English Language Arts, generated around a K-12 standards taxonomy for instruction-tuning and educational-content experiments. How this was built (read this first) These are programmatically generated training examples, not curriculum written by educators and not the text of any official standard. A generator combined standards metadata - codes, grade levels… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-ela-standards-expanded.texttext-generation10K<n<100K0 likes46 downloads2mo agoHugging Face30robworks-software /k12-special-education-accommodations K-12 Special Education Accommodations A small reference dataset of special education accommodations and the federal IDEA disability taxonomy. This is a reference table, not a corpus - 50 accommodation records plus two small lookup tables. Loading from datasets import load_dataset ds = load_dataset("robworks-software/k12-special-education-accommodations") Contents Table Rows Contents train / validation / test 40 / 5 / 5 accommodation… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-special-education-accommodations.texttext-classificationn<1K0 likes46 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.