datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aimo3-math-dataset
AIMO3 Math Dataset
Training data for AI Mathematical Olympiad Progress Prize 3.
Files
train_cot.jsonl - Chain-of-Thought examples
train_tir.jsonl - Tool-Integrated Reasoning examples
Author
Ryan J Cardwell (Archer Phoenix) - AIMO3 Competitor
AI-MO-NuminaMath-TIR-korean-240918
IMPORTANT NOTE
This data is part of the progress. Current translation progress: 24.85% (2024-09-18 01:32 KST)
I'm taking a short break due to personal reasons. I'll be back in a month.
TODO-LIST
Finish translation
Translation
I used gemini-1.5-pro-exp-0827. The prompt used for translation will be disclosed at the end.
Dataset Card for NuminaMath CoT
Dataset Summary
Tool-integrated reasoning (TIR) plays a crucial role in this… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/AI-MO-NuminaMath-TIR-korean-240918.moh_8_fake_rollouts
MOH-8 Fake Rollouts
480 math competition problems, each with 8 candidate solution rollouts from OSS 120B.
A controlled number of rollouts per problem are correct — use this to train/test a
verifier model that must identify which solutions are right.
Source
Problems and rollouts sampled from aimosprite/training-data-oss120b (the oss128-fixed-FINAL.jsonl file).
Only polymath-source problems in the 2/8–4/8 pass rate range (32–64 correct out of 128 attempts).
4 problems… See the full description on the dataset page: https://huggingface.co/datasets/aimosprite/moh_8_fake_rollouts.
