datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
competition_math
Dataset Card for Mathematics Aptitude Test of Heuristics (MATH) dataset
Dataset Summary
The Mathematics Aptitude Test of Heuristics (MATH) dataset consists of problems
from mathematics competitions, including the AMC 10, AMC 12, AIME, and more.
Each problem in MATH has a full step-by-step solution, which can be used to teach
models to generate answer derivations and explanations.
Supported Tasks and Leaderboards
[More Information Needed]
Languages… See the full description on the dataset page: https://huggingface.co/datasets/qwedsacf/competition_math.competition_mathcompetition_math
Dataset Card for "competition_math"
More Information needed
competition_math_hf_dataset
Dataset Card for "competition_math_hf_dataset"
Homepage - https://huggingface.co/datasets/hendrycks/competition_math
This is just the competetion math dataset, put in HF dataset format for ease of use with any finetuning tasks
competition_math_imagessft-ready-hendrycks-competition_mathcompetition_math
Dataset Card for "competition_math"
Added column with final solution extracted from \boxed{} tags.
Added numeric congig that only contains questions with numeric answers.
Dataset Summary
MATH contains 12,500 challenging competition mathematics problems. Each problem in MATH has a full step-by-step solution which can be used to teach models to generate answer derivations and explanation
This dataset card aims to be a base template for new datasets.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/jeggers/competition_math.highschool_math_competition初高中 中文数学竞赛pdf书籍。
deepseek_ocr.zip是使用https://github.com/deepseek-ai/DeepSeek-OCR/得到的OCR文本
forum-competition-math-training-pool
Forum competition mathematics training pool
Olympiad and contest mathematics from three public datasets, gathered at pinned revisions and
shipped twice over. sources/ holds each dataset the way its publisher ships it, in its own file
format with its own fields and nothing renamed, 287091 rows across three folders. pool/ holds the
union of those same datasets in one format, one JSON object per line, deduplicated by problem text
and reduced to 282140 rows, every row labelled with… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/forum-competition-math-training-pool.competition-math-training-pool
Competition mathematics training pool
Public competition mathematics, six datasets gathered at pinned revisions, shipped twice over.
sources/ holds each dataset the way its publisher ships it, in its own file format with its own
fields and nothing renamed, 1951046 rows across six folders. pool/ holds the union of those
same datasets in one format, one JSON object per line, deduplicated by problem text and reduced
to 1125451 rows, every row labelled with the dataset it came from… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/competition-math-training-pool.competition_math_selectedMaths_competition_questionsOpenMathInstruct-2-augmented-mathnvidia/OpenMathInstruct-2のaugmented_mathだけ抜き出したものです.licenseは元データと同じです.
math-competition-evalhendrycks_competition_math_N_A
hendrycks_competition_math
Dataset Description
This dataset contains evaluation results for hendrycks_competition_math with label column N_A, with various model performance metrics and samples.
Dataset Summary
The dataset contains original samples from the evaluation process, along with metadata like model names, input columns, and scores. This helps with understanding model performance across different tasks and datasets.
Features
id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/gallifantjack/hendrycks_competition_math_N_A.competition-answer-math-training-pool
Competition answer mathematics training pool
Competition mathematics problems that ask for a single final answer, three public datasets gathered
at pinned revisions, shipped twice over. sources/ holds each dataset the way its publisher ships
it, in its own file format with its own fields and nothing renamed, 113045 rows across three
folders. pool/ holds the union of those same datasets in one format, one JSON object per line,
deduplicated by problem text and reduced to 107637… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/competition-answer-math-training-pool.Competition-level_Mathematics_Physics_Reasoning_Corpus
Title
Competition-level Mathematics, Physics Reasoning Corpus
Size
50.000+ text + multimodal competition level problems, each with step-by-step solutions and final answers
Format
Natural language explanations with multimodal samples include images (graphs, diagrams, etc.)
Subject
Mathematics/Physics and etc
Labeling Details
Question ID/Question Stem (Full text/content) /Subject/Question Type (Multiple Choice/Short Answer format… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/Competition-level_Mathematics_Physics_Reasoning_Corpus.John_O_Bryan_Mathematics_Competitioncompetition-math-japanese
Competition Math (Japanese Translation)
qwedsacf/competition_math の日本語翻訳版です。
データセットの説明
数学コンペティションレベルの問題、解答、解法を含むデータセットです。元のデータセット(MATH dataset)を日本語に翻訳しました。
難易度はLevel 1〜Level 5の5段階、分野は7種類(Algebra, Counting & Probability, Geometry, Intermediate Algebra, Number Theory, Prealgebra, Precalculus)に分類されています。
使用方法
from datasets import load_dataset
ds = load_dataset("kfsky/competition-math-japanese")
print(ds["train"][0])
カラム
カラム名
説明
problem… See the full description on the dataset page: https://huggingface.co/datasets/kfsky/competition-math-japanese.competition_math_dspyqwen3-8b-high-school-math-competition-depth2-val3competition_math_llama3.2competition_math_stratified
Notes
Source dataset: qwedsacf/competition_math.
level is cast to a ClassLabel and used for stratified splitting.
Splits are re-created with seed 42: 90% train / 5% validation / 5% test, stratified by level.
Final answer is extracted from the last \boxed{...} in solution; empty/invalid extractions are dropped.
Validation checks ensure every sample has an answer and that it appears in the original solution (case-insensitive).
Dropped (no extractable answer): {'train': 4… See the full description on the dataset page: https://huggingface.co/datasets/imyangyixuan/competition_math_stratified.competition_math
Dataset Card for Mathematics Aptitude Test of Heuristics (MATH) dataset
Dataset Summary
The Mathematics Aptitude Test of Heuristics (MATH) dataset consists of problems
from mathematics competitions, including the AMC 10, AMC 12, AIME, and more.
Each problem in MATH has a full step-by-step solution, which can be used to teach
models to generate answer derivations and explanations.
Supported Tasks and Leaderboards
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/Seonwan/competition_math.competition_math
Dataset Card for Mathematics Aptitude Test of Heuristics (MATH) dataset
Dataset Summary
The Mathematics Aptitude Test of Heuristics (MATH) dataset consists of problems
from mathematics competitions, including the AMC 10, AMC 12, AIME, and more.
Each problem in MATH has a full step-by-step solution, which can be used to teach
models to generate answer derivations and explanations.
Supported Tasks and Leaderboards
[More Information Needed]
Languages… See the full description on the dataset page: https://huggingface.co/datasets/lucasvincent/competition_math.competition_mathMAIR-Competition-Math__mtebcompetition_math_with_retrievalTurkce-hendrycks_competition_math
This dataset is a machine-translated version of lighteval/MATH-Hard. We translated it using machine translation for the Teknofest 2024 Natural Language Processing competition.
competition_math-level_5
