datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
orca-math-word-problems-200k
Dataset Card
This dataset contains ~200K grade school math word problems. All the answers in this dataset is generated using Azure GPT4-Turbo. Please refer to Orca-Math: Unlocking the potential of
SLMs in Grade School Math for details about the dataset construction.
Dataset Sources
Repository: microsoft/orca-math-word-problems-200k
Paper: Orca-Math: Unlocking the potential of
SLMs in Grade School Math
Direct Use
This dataset has been designed to… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/orca-math-word-problems-200k.orca-math-word-problems-200k
Dataset Card for Orca Math Word Problems 200k
This is a formatted version of microsoft/orca-math-word-problems-200k to store the conversations in the same format as the OpenAI SDK.
SFT-orca-math-word-problems-200korca-math-word-problems-193k-korean원본 데이터셋: https://huggingface.co/datasets/microsoft/orca-math-word-problems-200k
번역 모델: Seagull-13b-translation
후처리
번역 repetition 오류 제거
LaTeX 오류 체크(전부는 아닐 수 있음. /(/) -> /(/ 같은 오류 등...)
Citation
@misc{mitra2024orcamath,
title={Orca-Math: Unlocking the potential of SLMs in Grade School Math},
author={Arindam Mitra and Hamed Khanpour and Corby Rosset and Ahmed Awadallah},
year={2024},
eprint={2402.14830},
archivePrefix={arXiv},
primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/kuotient/orca-math-word-problems-193k-korean.orca-math-word-problems-80kI removed samples where "question" character length was over 1,000 and "answer" character length was over 2,000, then randomly subsampled 80k rows.
orca-math-word-problems-trorca-math-word-problemsmlabonne-orca-math-word-problems-80korca-math-word-problems-200k-ruorca-math-word-problems-10002_20004orca-math-word-problems-140028_150030orca-math-word-problems-180036_190038orca-math-word-problems-30006_40008Vietnamese-microsoft-orca-math-word-problems-200k-gg-translatedorca-math-word-problems-90018_100020Elementary_Math_Word_Problems_LLM_Training_Short
Dataset Card for Math Problem Generator
Dataset Summary
This dataset contains a subset of 100,000 procedurally generated math word problems, covering various mathematical concepts and difficulty levels. The problems were generated using a Java program that creates contextual word problems with detailed solutions and explanations.
🔗 Full dataset available on Gumroad
The full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com
Want more?… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Elementary_Math_Word_Problems_LLM_Training_Short.orca-math-word-problems-40008_50010Math2Visual-Generating_Pedagogically_Meaningful_Visuals_for_Math_Word_Problems
📚 Dataset Overview
This dataset accompanies the ACL 2025 Findings paper:"Generating Pedagogically Meaningful Visuals for Math Word Problems: A New Benchmark and Analysis of Text-to-Image Models".
📄 ACL 2025 Findings Paper — Math2Visual
🎥 Project Video
🤖 Visual Language Generation Model
💻 GitHub Codebase
📦 Contents
final_annotated_visual_language_dataset_updated.csvContains a curated set of primary school math word problems along with their corresponding… See the full description on the dataset page: https://huggingface.co/datasets/junling24/Math2Visual-Generating_Pedagogically_Meaningful_Visuals_for_Math_Word_Problems.adaption-math-and-word-problem-solutions
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-math_and_word_problem_solutions
This dataset contains pairs of mathematical and arithmetic word problems alongside their detailed step-by-step solutions. The content ranges from elementary arithmetic scenarios to advanced competition-level proofs involving algebra, number theory, and combinatorics. Solutions often include intermediate calculations marked with specific formatting tags… See the full description on the dataset page: https://huggingface.co/datasets/Minutor/adaption-math-and-word-problem-solutions.orca-math-word-problems-70014_80016orca-math-word-problems-200k-askllm-v1
orca-math-word-problems-200k-askllm-v1
データセット microsoft/orca-math-word-problems-200k に対して、 Ask-LLM 手法でスコア付けしたデータセットです。
元データセットのカラムに加え askllm_score というカラムが追加されており、ここに Ask-LLM のスコアが格納されています。
Ask-LLM でスコア付けに使用した LLM は Rakuten/RakutenAI-7B-instruct で、プロンプトは以下の通りです。
###
{data}
###
Does the previous paragraph demarcated within ### and ### contain informative signal for pre-training a large-language model? An informative datapoint should be well-formatted, contain some usable knowledge of… See the full description on the dataset page: https://huggingface.co/datasets/geniacllm/orca-math-word-problems-200k-askllm-v1.orca-math-word-problems-170034_180036orca-math-word-problems-120024_130026orca-math-word-problems-200k-turkmen
Turkmen Orca Math Word Problems 200k Dataset
Overview
This dataset is a Turkmen translation of the original microsoft/orca-math-word-problems-200k dataset. The Orca Math Word Problems dataset contains 200,000 high-quality math word problems and their solutions. This Turkmen version aims to extend the accessibility of math problem-solving datasets to the Turkmen language community.
Dataset Details
Original Dataset: microsoft/orca-math-word-problems-200k… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/orca-math-word-problems-200k-turkmen.orca-math-word-problems-100k-en-zh-mix100k English and Chinese mixed version of microsoft/orca-math-word-problems-200k
orca-math-word-problems-20004_30006orca-math-word-problems-110022_120024orca-math-word-problems-10002_20004-spanishorca-math-word-problems-0_10002orca-math-word-problems-80016_90018
