CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aslawliet /cn-k12text100K<n<1M13 likes2.3k downloads2y agoHugging Face02alexkstern /dyck-k128-seq_len_2048-1B dyck-k128-seq_len_2048-1B Procedurally generated k-shuffle Dyck bracket sequences (Hu et al. 2025, arXiv:2502.19249), as flat uint16 token-id .bin files. Token ids are 0-based: opening bracket type i is id i and its matching close is i + k, so ids span [0, 2k) and the vocabulary is 2k = 256. Grammar parameters param value k (bracket types) 128 max_depth 16 p_open 0.5 seq_length 2048 file split tokens train.bin train 999,999,488 val.bin val 10,000… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/dyck-k128-seq_len_2048-1B.tabularn<1K0 likes1.5k downloads4mo agoHugging Face03k19862217 /simpsons_script_linestext10K<n<100K0 likes1.2k downloads3y agoHugging Face04tunaaa126 /K12-Dataset K12-KGraph K12-KGraph is a curriculum-aligned knowledge graph built from official People's Education Press (PEP) K-12 textbooks. It focuses on curriculum cognition, namely the structured understanding of how school knowledge is organized, connected, and sequenced. The current release covers mathematics, physics, chemistry, and biology across primary, middle, and high school, and includes three resources derived from the same graph: K12-KGraph: the core knowledge graph K12-Bench: a… See the full description on the dataset page: https://huggingface.co/datasets/tunaaa126/K12-Dataset.text1K<n<10K0 likes678 downloads5mo agoHugging Face05lfaviate /China-K12-STEM-10K-CoT-Reasoning K12-STEM-CoT-Chinese 1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams. The largest structured Chinese math/physics/chemistry reasoning dataset. This is a curated sample (10,000 problems) of the full 1.54M dataset available via API. Full Dataset Access Access the full 1,540,000+ problems via API → This Sample Full API Total problems 10,025 1,540,000+ With CoT solutions 10,025 1,490,000+ With diagrams 6,093 740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.tabularquestion-answering10K<n<100K3 likes604 downloads7mo agoHugging Face06marin-community /openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16 OpenThoughts-4 Code SDG: Qwen3-30B-A3B-Thinking-2507 (n=16, top-16 logprobs) Synthetic generations from Qwen/Qwen3-30B-A3B-Thinking-2507 on the Marin OpenThoughts-4 code SDG prompt set. Each prompt is sampled n=16 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes465 downloads5mo agoHugging Face07k1000dai /libero-pickandplace-segment-next-scene-ab-2image100K<n<1M0 likes457 downloads9mo agoHugging Face08marin-community /openthoughts4-code-9168-prompts-qwen3-32b-n16-flattened-logprobs-k16 OpenThoughts-4 Code SDG: Qwen3-32B (n=16, top-16 logprobs) Synthetic generations from Qwen/Qwen3-32B on the Marin OpenThoughts-4 code SDG prompt set. Each prompt is sampled n=16 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-32b-n16-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes445 downloads5mo agoHugging Face09tangchen-ai /birdcode-deepswe-k1d BirdCode on DeepSWE — k=1, single attempt, no web tools ⚠️ Reading the metric correctly: The summary card's "Average f2p 0.89" is the test-case-level pass fraction (f2p_passed/f2p_total, averaged per task) — it is NOT the official DeepSWE leaderboard metric. The official binary score is the reward field (1 only when all F2P and P2P tests pass): 60/113 = 0.531. Per-trial reward values are visible in each trial's rewards block below. Evaluation of BirdCode (a from-scratch… See the full description on the dataset page: https://huggingface.co/datasets/tangchen-ai/birdcode-deepswe-k1d.tabularn<1K0 likes419 downloads10d agoHugging Face10zhenliuu /k12-multidisciplinary K12 多学科图文推理数据集 面向中小学数学、物理、生物、地理和化学的多学科图文推理数据。 GitHub 主页与训练代码 数据范围 配置 split 题目数 含图题数 唯一图片数 default raw 735,650 515,089 416,099 dapo train 1,160 712 719 原始数据与训练集按用途分别提供。训练集从总数据集中筛选整理而来,面向数学与物理推理任务,可用于不同模型与训练框架。 原始数据 数据由公开开源数据筛选、自建实体书OCR抽取及基于vLLM的合成与改写三部分构成,经过图片回收与校验、结构统一、来源标签清理、重复题处理及答案冲突复核。数据包含纯文本题和含图题,同时覆盖选择题与非选择题。 数学62,966题、物理164,398题、生物199,084题、地理174,282题、化学134,920题。含图题占70.02%。 数学与物理训练集… See the full description on the dataset page: https://huggingface.co/datasets/zhenliuu/k12-multidisciplinary.documentvisual-question-answering100K<n<1M0 likes233 downloads5d agoHugging Face11marin-community /openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-30B-A3B-Thinking-2507 (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-30B-A3B-Thinking-2507 on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes227 downloads5mo agoHugging Face12marin-community /openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-32B (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-32B on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes212 downloads5mo agoHugging Face13Neelectric /OpenR1-Math-cn_k12-91ktext10K<n<100K0 likes203 downloads1y agoHugging Face14xyliu6 /k12-freeformimage10K<n<100K2 likes186 downloads1y agoHugging Face15marin-community /openthoughts4-code-9168-prompts-qwen3-4b-n16-flattened-logprobs-k16 OpenThoughts-4 Code SDG: Qwen3-4B (n=16, top-16 logprobs) Synthetic generations from Qwen/Qwen3-4B on the Marin OpenThoughts-4 code SDG prompt set. Each prompt is sampled n=16 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model Qwen/Qwen3-4B… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-4b-n16-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes161 downloads5mo agoHugging Face16k1000dai /merged-libero-pickandplace-segment-v2-nohistoryimage100K<n<1M0 likes153 downloads8mo agoHugging Face17robworks-software /k12-standards-instruction-tasks K-12 Curriculum Tasks (generated) 2,489 generated instruction/input/output records covering five curriculum tasks: assessment creation, learning objective generation, misconception detection, standard explanation, and standards Q&A. Content is predominantly mathematics. Important: the name is misleading Despite the name, this dataset contains no school directory data. There are four columns - task, input, output, metadata - and no staff, principal, or school… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-standards-instruction-tasks.texttext-generation1K<n<10K1 likes150 downloads2mo agoHugging Face18robworks-software /us-k12-schools-directory US K-12 Schools Directory A directory of 124,613 US K-12 schools covering all 50 states, DC, and US territories, compiled from federal and state government sources. Each record carries directory information (address, phone, website), enrollment and demographics, and, where a source supplied it, a principal name and email. This is a compilation of public government data. It is not a survey, and no field was independently verified against the school itself. Loading… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/us-k12-schools-directory.tabulartabular-classification100K<n<1M0 likes150 downloads2mo agoHugging Face19k1seul /rr_three_tasks_v1 rr_three_tasks_v1 Three tasks on a Trossen AI solo arm, merged into one LeRobot v2.1 dataset. task episodes frames pick_specific_item_from_clutter 243 59088 pick_two_in_order 99 40478 open_pot_and_place 100 47288 meta/sources.jsonl maps every episode to its source dataset, episode and revision, with the staging record (open_pot_and_place variant, pick_two second object, sheet row). Held-out evaluation episodes meta/eval_episodes_v1.json: 44… See the full description on the dataset page: https://huggingface.co/datasets/k1seul/rr_three_tasks_v1.tabularroboticsn<1K0 likes145 downloads11d agoHugging Face20k1000dai /converted_mixed_pickandplace_datasetimage100K<n<1M0 likes128 downloads9mo agoHugging Face21k1000dai /libero-expert-selectionimage100K<n<1M0 likes118 downloads9mo agoHugging Face22k1000dai /merged-libero-pickandplace-segment-v2image100K<n<1M0 likes116 downloads8mo agoHugging Face23FoundryAILabs /k12-indian-curriculum-4.9m BharatLLM K-12 Indian Curriculum Dataset (4.9M) 4,904,936 question-answer pairs covering CBSE/NCERT K-12 curriculum across 12 Indian languages. Language Script Entries English Latin ~594K Hindi Devanagari ~449K Bengali Bengali ~408K Telugu Telugu ~408K Tamil Tamil ~408K Kannada Kannada ~408K Malayalam Malayalam ~408K Marathi Devanagari ~408K Gujarati Gujarati ~408K Odia Odia ~408K PunjabiGurmukhi ~408K Urdu Nastaliq ~374K Format {… See the full description on the dataset page: https://huggingface.co/datasets/FoundryAILabs/k12-indian-curriculum-4.9m.textquestion-answering1M<n<10M1 likes114 downloads6mo agoHugging Face24lab-flair /qa-dataset-k1000 QA Dataset K1000 — The First Drop of Ink Question-answering data with gold documents and distractor pools for long-context evaluation, accompanying The First Drop of Ink: Nonlinear Impact of Distracting Information in Long-Context Reasoning by Muhan Gao, Zih-Ching Chen, and Kuan-Hao Huang (ICML 2026). Paper · Full text (v2) · Hugging Face paper page The paper studies how the proportion of hard distractors affects performance at fixed context length. It reports a nonlinear… See the full description on the dataset page: https://huggingface.co/datasets/lab-flair/qa-dataset-k1000.textquestion-answeringn<1K1 likes107 downloads2d agoHugging Face25cygu /sampling-distill-train-data-kgw-k1-gamma0.25-delta1 Dataset Card for "sampling-distill-train-data-kgw-k1-gamma0.25-delta1" Training data for sampling-based watermark distillation using the KGW k=1,γ=0.25,δ=1k=1, \gamma=0.25, \delta=1k=1,γ=0.25,δ=1 watermarking strategy in the paper On the Learnability of Watermarks for Language Models. Llama 2 7B with decoding-based watermarking was used to generate 640,000 watermarked samples, each 256 tokens long. Each sample is prompted with 50-token prefixes from OpenWebText (prompts not included… See the full description on the dataset page: https://huggingface.co/datasets/cygu/sampling-distill-train-data-kgw-k1-gamma0.25-delta1.text100K<n<1M0 likes104 downloads2y agoHugging Face26k1000dai /libero-pickandplace-segment-expert-selectionimage100K<n<1M0 likes104 downloads9mo agoHugging Face27novastar111 /pacman_hard_cot_chunk_k10_train pacman_hard_cot_chunk_k10_train BAGEL VLM-Gym world-model dataset (pacman / cot). CoT chunk-K train set: all-step interleaved imagined reasoning; re-grounds on the true frame every K=10 steps. layout: Train-only. Gzipped-JSONL shards under training/; each row is one packed SFT sample with base64-JPEG frames inline. images are base64-encoded JPEG frames stored inline in each JSONL row. Pairs with the matching pacman checkpoint(s) under the companion model org; CoT and non-CoT… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/pacman_hard_cot_chunk_k10_train.tabular100K<n<1M0 likes101 downloads1mo agoHugging Face28krittus /k12-indian-curriculum-4.9m BharatLLM K-12 Indian Curriculum Dataset (4.9M) 4,904,936 question-answer pairs covering CBSE/NCERT K-12 curriculum across 12 Indian languages. Language Script Entries English Latin ~594K Hindi Devanagari ~449K Bengali Bengali ~408K Telugu Telugu ~408K Tamil Tamil ~408K Kannada Kannada ~408K Malayalam Malayalam ~408K Marathi Devanagari ~408K Gujarati Gujarati ~408K Odia Odia ~408K Punjabi Gurmukhi ~408K Urdu Nastaliq ~374K Format… See the full description on the dataset page: https://huggingface.co/datasets/krittus/k12-indian-curriculum-4.9m.textquestion-answering1M<n<10M1 likes94 downloads6mo agoHugging Face29opendatalab /K12textbook覆盖小学、初中、高中的高质量中文K12教材语料,经过精细的文本抽取和数据处理,可用于学术研究 texttext-generationn<1K3 likes88 downloads1y agoHugging Face30SchoolData /us-k12-schools US K–12 Schools — Open Dataset for the AI Era One encyclopedic paragraph for every one of the 122,675 K–12 schools in the United States, ready to use as pretraining text. Ask a language model about a large suburban high school and it will answer. Ask it about the K–8 school in a rural county of 4,000 people and it has nothing to say — because nothing about that school was ever written down on the open web. Only about 12% of American schools have a Wikipedia article at all, and… See the full description on the dataset page: https://huggingface.co/datasets/SchoolData/us-k12-schools.texttext-generation100K<n<1M0 likes84 downloads19d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.