datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PolyMath
Dataset Card for PolyMath
Dataset Summary
PolyMath is a curated dataset of 11,090 high-difficulty mathematical problems designed for training reasoning models. Built for the AIMO Math Corpus Prize. Existing math datasets (NuminaMath-1.5, OpenMathReasoning) suffer from high noise rates in their hardest samples and largely unusable proof-based problems.
PolyMath addresses both issues through:
Data scraping: problems sourced from official competition PDFs absent from… See the full description on the dataset page: https://huggingface.co/datasets/AIMO-Corpus/PolyMath.COPSD-PolyMath-TrainDataset
Crosslingual On-Policy Self-Distillation for Multilingual Reasoning
This repository contains the dataset for the paper Crosslingual On-Policy Self-Distillation for Multilingual Reasoning.
The project proposes Crosslingual On-Policy Self-Distillation (COPSD), a method that transfers a model's high-resource reasoning behavior to low-resource languages by using the model as both a student and a teacher with privileged crosslingual context.
Links
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/yihongLiu/COPSD-PolyMath-TrainDataset.polymath-reasoning-v1
Polymath Reasoning Corpus v1
A multi-domain reasoning corpus that pairs explicit <thinking> chain-of-thought
with substantive output across eleven distinct cognitive formats — chain-of-thought
debates, multi-figure salons, counterfactual perspective pieces, logic problems
with worked solutions, multi-step reasoning chains, math/science/code explanations,
and challenging Q&A.
The unifying angle is how experts actually reason across disciplines, not just what they
conclude. Most CoT… See the full description on the dataset page: https://huggingface.co/datasets/MonumentalSystems/polymath-reasoning-v1.
