datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mathlib-Normalized-Sexpr
Mathlib Normalized S-Expressions
Lean 4 proof states from Mathlib, paired with the tactic applied at each
step, in three representations extracted directly from the Lean kernel:
Source-faithful S-expressions of the goal and every hypothesis, as
Lean elaborated them.
Normalized S-expressions of the same state, with stable local-context
indices suitable for model input.
Annotated tactic syntax -- the original tactic's syntax tree with
identifier leaves resolved to the constants… See the full description on the dataset page: https://huggingface.co/datasets/jajostrains/Mathlib-Normalized-Sexpr.mathlib4-state-changemathlib-informal-splitmathlib_RL_v3canonical-drafter-extract-mathlib
Canonical Drafter — extract data (raw)
Ground-truth drafter/premise data lifted from existing Mathlib proofs by
Training/ExtractData.lean. This is the raw pool (429303 draft rows,
664445 premise rows from: Mathlib). Every row is a real have / closed
subgoal taken from a checked proof, so there are no success/used flags to filter
on. Columns follow the uniform schema shared with the rollout dataset, so the two
pools concatenate cleanly.
config drafts
One row per… See the full description on the dataset page: https://huggingface.co/datasets/awhecmu/canonical-drafter-extract-mathlib.mathlib_v2mathlib_RL_v3_goalsmathlib_benchmark_v1mathlib_RL_v1mathlib_RL_v2mathlib_RL_v3_sortedmathlib_RL_v3_meta_tactic_3mathlib_RL_v3_iter11mathlib_RL_length_bracketsmathlib_RL_v3_tracedmathlib_v15mathlib_v09mathlib_benchmark_v09_2048mathlib_RL_v3_lengthmathlib_benchmarkmathlib_v3mathlib_benchmark_v09mathlib_RL_v4Mathlib_RL_V13mathlib_RL_exp_lengthmathlib_RL_v3_traced1mathlib_RL_v3_traced2mathlib_benchmark_v15mathlib_benchmark_v09_newmathlib_RL_eval_complexity
