datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cn-k12dyck-k128-seq_len_2048-1B
dyck-k128-seq_len_2048-1B
Procedurally generated k-shuffle Dyck bracket sequences (Hu et al. 2025, arXiv:2502.19249), as flat uint16 token-id .bin files. Token ids are 0-based: opening bracket type i is id i and its matching close is i + k, so ids span [0, 2k) and the vocabulary is 2k = 256.
Grammar parameters
param
value
k (bracket types)
128
max_depth
16
p_open
0.5
seq_length
2048
file
split
tokens
train.bin
train
999,999,488
val.bin
val
10,000… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/dyck-k128-seq_len_2048-1B.simpsons_script_linesK12-Dataset
K12-KGraph
K12-KGraph is a curriculum-aligned knowledge graph built from official People's Education Press (PEP) K-12 textbooks. It focuses on curriculum cognition, namely the structured understanding of how school knowledge is organized, connected, and sequenced.
The current release covers mathematics, physics, chemistry, and biology across primary, middle, and high school, and includes three resources derived from the same graph:
K12-KGraph: the core knowledge graph
K12-Bench: a… See the full description on the dataset page: https://huggingface.co/datasets/tunaaa126/K12-Dataset.China-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16
OpenThoughts-4 Code SDG: Qwen3-30B-A3B-Thinking-2507 (n=16, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-30B-A3B-Thinking-2507
on the Marin OpenThoughts-4 code SDG prompt
set.
Each prompt is sampled n=16 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16.libero-pickandplace-segment-next-scene-ab-2openthoughts4-code-9168-prompts-qwen3-32b-n16-flattened-logprobs-k16
OpenThoughts-4 Code SDG: Qwen3-32B (n=16, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-32B
on the Marin OpenThoughts-4 code SDG prompt
set.
Each prompt is sampled n=16 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field
Value
Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-32b-n16-flattened-logprobs-k16.birdcode-deepswe-k1d
BirdCode on DeepSWE — k=1, single attempt, no web tools
⚠️ Reading the metric correctly: The summary card's "Average f2p 0.89" is the
test-case-level pass fraction (f2p_passed/f2p_total, averaged per task) — it is
NOT the official DeepSWE leaderboard metric. The official binary score is the
reward field (1 only when all F2P and P2P tests pass): 60/113 = 0.531.
Per-trial reward values are visible in each trial's rewards block below.
Evaluation of BirdCode (a from-scratch… See the full description on the dataset page: https://huggingface.co/datasets/tangchen-ai/birdcode-deepswe-k1d.k12-multidisciplinary
K12 多学科图文推理数据集
面向中小学数学、物理、生物、地理和化学的多学科图文推理数据。
GitHub 主页与训练代码
数据范围
配置
split
题目数
含图题数
唯一图片数
default
raw
735,650
515,089
416,099
dapo
train
1,160
712
719
原始数据与训练集按用途分别提供。训练集从总数据集中筛选整理而来,面向数学与物理推理任务,可用于不同模型与训练框架。
原始数据
数据由公开开源数据筛选、自建实体书OCR抽取及基于vLLM的合成与改写三部分构成,经过图片回收与校验、结构统一、来源标签清理、重复题处理及答案冲突复核。数据包含纯文本题和含图题,同时覆盖选择题与非选择题。
数学62,966题、物理164,398题、生物199,084题、地理174,282题、化学134,920题。含图题占70.02%。
数学与物理训练集… See the full description on the dataset page: https://huggingface.co/datasets/zhenliuu/k12-multidisciplinary.openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16
OpenThoughts-4 Science SDG: Qwen3-30B-A3B-Thinking-2507 (n=8, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-30B-A3B-Thinking-2507
on the Marin OpenThoughts-4 science SDG prompt
set.
Each prompt is sampled n=8 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16.openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16
OpenThoughts-4 Science SDG: Qwen3-32B (n=8, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-32B
on the Marin OpenThoughts-4 science SDG prompt
set.
Each prompt is sampled n=8 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field
Value
Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16.OpenR1-Math-cn_k12-91kk12-freeformopenthoughts4-code-9168-prompts-qwen3-4b-n16-flattened-logprobs-k16
OpenThoughts-4 Code SDG: Qwen3-4B (n=16, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-4B
on the Marin OpenThoughts-4 code SDG prompt
set.
Each prompt is sampled n=16 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field
Value
Generator model
Qwen/Qwen3-4B… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-4b-n16-flattened-logprobs-k16.merged-libero-pickandplace-segment-v2-nohistoryk12-standards-instruction-tasks
K-12 Curriculum Tasks (generated)
2,489 generated instruction/input/output records covering five curriculum tasks:
assessment creation, learning objective generation, misconception detection, standard
explanation, and standards Q&A. Content is predominantly mathematics.
Important: the name is misleading
Despite the name, this dataset contains no school directory data. There are four
columns - task, input, output, metadata - and no staff, principal, or school… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-standards-instruction-tasks.us-k12-schools-directory
US K-12 Schools Directory
A directory of 124,613 US K-12 schools covering all 50 states, DC, and US territories,
compiled from federal and state government sources. Each record carries directory
information (address, phone, website), enrollment and demographics, and, where a source
supplied it, a principal name and email.
This is a compilation of public government data. It is not a survey, and no field was
independently verified against the school itself.
Loading… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/us-k12-schools-directory.rr_three_tasks_v1
rr_three_tasks_v1
Three tasks on a Trossen AI solo arm, merged into one LeRobot v2.1 dataset.
task
episodes
frames
pick_specific_item_from_clutter
243
59088
pick_two_in_order
99
40478
open_pot_and_place
100
47288
meta/sources.jsonl maps every episode to its source dataset, episode and revision, with the
staging record (open_pot_and_place variant, pick_two second object, sheet row).
Held-out evaluation episodes
meta/eval_episodes_v1.json: 44… See the full description on the dataset page: https://huggingface.co/datasets/k1seul/rr_three_tasks_v1.converted_mixed_pickandplace_datasetlibero-expert-selectionmerged-libero-pickandplace-segment-v2k12-indian-curriculum-4.9m
BharatLLM K-12 Indian Curriculum Dataset (4.9M)
4,904,936 question-answer pairs covering CBSE/NCERT K-12 curriculum across 12 Indian languages.
Language
Script
Entries
English
Latin
~594K
Hindi
Devanagari
~449K
Bengali
Bengali
~408K
Telugu
Telugu
~408K
Tamil
Tamil
~408K
Kannada
Kannada
~408K
Malayalam
Malayalam
~408K
Marathi
Devanagari
~408K
Gujarati
Gujarati
~408K
Odia
Odia
~408K
PunjabiGurmukhi
~408K
Urdu
Nastaliq
~374K
Format
{… See the full description on the dataset page: https://huggingface.co/datasets/FoundryAILabs/k12-indian-curriculum-4.9m.qa-dataset-k1000
QA Dataset K1000 — The First Drop of Ink
Question-answering data with gold documents and distractor pools for long-context evaluation, accompanying The First Drop of Ink: Nonlinear Impact of Distracting Information in Long-Context Reasoning by Muhan Gao, Zih-Ching Chen, and Kuan-Hao Huang (ICML 2026).
Paper · Full text (v2) · Hugging Face paper page
The paper studies how the proportion of hard distractors affects performance at fixed context length. It reports a nonlinear… See the full description on the dataset page: https://huggingface.co/datasets/lab-flair/qa-dataset-k1000.sampling-distill-train-data-kgw-k1-gamma0.25-delta1
Dataset Card for "sampling-distill-train-data-kgw-k1-gamma0.25-delta1"
Training data for sampling-based watermark distillation using the KGW k=1,γ=0.25,δ=1k=1, \gamma=0.25, \delta=1k=1,γ=0.25,δ=1 watermarking strategy in the paper On the Learnability of Watermarks for Language Models. Llama 2 7B with decoding-based watermarking was used to generate 640,000 watermarked samples, each 256 tokens long. Each sample is prompted with 50-token prefixes from OpenWebText (prompts not included… See the full description on the dataset page: https://huggingface.co/datasets/cygu/sampling-distill-train-data-kgw-k1-gamma0.25-delta1.libero-pickandplace-segment-expert-selectionpacman_hard_cot_chunk_k10_train
pacman_hard_cot_chunk_k10_train
BAGEL VLM-Gym world-model dataset (pacman / cot).
CoT chunk-K train set: all-step interleaved imagined reasoning; re-grounds on the true frame every K=10 steps.
layout: Train-only. Gzipped-JSONL shards under training/; each row is one packed SFT sample with base64-JPEG frames inline.
images are base64-encoded JPEG frames stored inline in each JSONL row.
Pairs with the matching pacman checkpoint(s) under the companion model org; CoT and non-CoT… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/pacman_hard_cot_chunk_k10_train.k12-indian-curriculum-4.9m
BharatLLM K-12 Indian Curriculum Dataset (4.9M)
4,904,936 question-answer pairs covering CBSE/NCERT K-12 curriculum across 12 Indian languages.
Language
Script
Entries
English
Latin
~594K
Hindi
Devanagari
~449K
Bengali
Bengali
~408K
Telugu
Telugu
~408K
Tamil
Tamil
~408K
Kannada
Kannada
~408K
Malayalam
Malayalam
~408K
Marathi
Devanagari
~408K
Gujarati
Gujarati
~408K
Odia
Odia
~408K
Punjabi
Gurmukhi
~408K
Urdu
Nastaliq
~374K
Format… See the full description on the dataset page: https://huggingface.co/datasets/krittus/k12-indian-curriculum-4.9m.K12textbook覆盖小学、初中、高中的高质量中文K12教材语料,经过精细的文本抽取和数据处理,可用于学术研究
us-k12-schools
US K–12 Schools — Open Dataset for the AI Era
One encyclopedic paragraph for every one of the 122,675 K–12 schools in the United
States, ready to use as pretraining text.
Ask a language model about a large suburban high school and it will answer. Ask it about
the K–8 school in a rural county of 4,000 people and it has nothing to say — because
nothing about that school was ever written down on the open web. Only about 12% of
American schools have a Wikipedia article at all, and… See the full description on the dataset page: https://huggingface.co/datasets/SchoolData/us-k12-schools.
