datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
orca-math-word-problems-200k
Dataset Card
This dataset contains ~200K grade school math word problems. All the answers in this dataset is generated using Azure GPT4-Turbo. Please refer to Orca-Math: Unlocking the potential of
SLMs in Grade School Math for details about the dataset construction.
Dataset Sources
Repository: microsoft/orca-math-word-problems-200k
Paper: Orca-Math: Unlocking the potential of
SLMs in Grade School Math
Direct Use
This dataset has been designed to… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/orca-math-word-problems-200k.orca-math-word-problems-200k
Dataset Card for Orca Math Word Problems 200k
This is a formatted version of microsoft/orca-math-word-problems-200k to store the conversations in the same format as the OpenAI SDK.
SFT-orca-math-word-problems-200korca-math-word-problems-193k-korean원본 데이터셋: https://huggingface.co/datasets/microsoft/orca-math-word-problems-200k
번역 모델: Seagull-13b-translation
후처리
번역 repetition 오류 제거
LaTeX 오류 체크(전부는 아닐 수 있음. /(/) -> /(/ 같은 오류 등...)
Citation
@misc{mitra2024orcamath,
title={Orca-Math: Unlocking the potential of SLMs in Grade School Math},
author={Arindam Mitra and Hamed Khanpour and Corby Rosset and Ahmed Awadallah},
year={2024},
eprint={2402.14830},
archivePrefix={arXiv},
primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/kuotient/orca-math-word-problems-193k-korean.orca-math-word-problems-trorca-math-word-problems-80kI removed samples where "question" character length was over 1,000 and "answer" character length was over 2,000, then randomly subsampled 80k rows.
orca-math-word-problemsmlabonne-orca-math-word-problems-80kdanish-word-problems-v2orca-math-word-problems-200k-ruorca-math-word-problems-193k-korean-jsonl원본 데이터셋
https://huggingface.co/datasets/microsoft/orca-math-word-problems-200k
https://huggingface.co/datasets/kuotient/orca-math-word-problems-193k-korean
Citation
@misc{mitra2024orcamath,
title={Orca-Math: Unlocking the potential of SLMs in Grade School Math},
author={Arindam Mitra and Hamed Khanpour and Corby Rosset and Ahmed Awadallah},
year={2024},
eprint={2402.14830},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
orca-math-word-problems-10002_20004adaption-arithmetic-algebra-word-problems
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-arithmetic-algebra-word-problems
This dataset features instruction and response pairs containing grade-school arithmetic and algebra word problems paired with step-by-step solutions. Problems cover multi-step arithmetic, percentages, ratios, linear equations, and simple systems solvable in two to five steps. Each completion demonstrates explicit reasoning and concludes with a… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/adaption-arithmetic-algebra-word-problems.orca-math-word-problems-30006_40008orca-math-word-problems-140028_150030orca-math-word-problems-180036_190038orca-math-word-problems-90018_100020lemonseed-word-problems
lemonseed-word-problems
LemonSeed — 13-category arithmetic word problems with plan + scratchpad.
Contents
wordproblems.jsonl (2500 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
Elementary_Math_Word_Problems_LLM_Training_Short
Dataset Card for Math Problem Generator
Dataset Summary
This dataset contains a subset of 100,000 procedurally generated math word problems, covering various mathematical concepts and difficulty levels. The problems were generated using a Java program that creates contextual word problems with detailed solutions and explanations.
🔗 Full dataset available on Gumroad
The full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com
Want more?… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Elementary_Math_Word_Problems_LLM_Training_Short.orca-math-word-problems-40008_50010orca-math-word-problems-200k-askllm-v1
orca-math-word-problems-200k-askllm-v1
データセット microsoft/orca-math-word-problems-200k に対して、 Ask-LLM 手法でスコア付けしたデータセットです。
元データセットのカラムに加え askllm_score というカラムが追加されており、ここに Ask-LLM のスコアが格納されています。
Ask-LLM でスコア付けに使用した LLM は Rakuten/RakutenAI-7B-instruct で、プロンプトは以下の通りです。
###
{data}
###
Does the previous paragraph demarcated within ### and ### contain informative signal for pre-training a large-language model? An informative datapoint should be well-formatted, contain some usable knowledge of… See the full description on the dataset page: https://huggingface.co/datasets/geniacllm/orca-math-word-problems-200k-askllm-v1.orca-math-word-problems-10002_20004-spanishorca-math-word-problems-170034_180036orca-math-word-problems-200k-turkmen
Turkmen Orca Math Word Problems 200k Dataset
Overview
This dataset is a Turkmen translation of the original microsoft/orca-math-word-problems-200k dataset. The Orca Math Word Problems dataset contains 200,000 high-quality math word problems and their solutions. This Turkmen version aims to extend the accessibility of math problem-solving datasets to the Turkmen language community.
Dataset Details
Original Dataset: microsoft/orca-math-word-problems-200k… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/orca-math-word-problems-200k-turkmen.Vietnamese-microsoft-orca-math-word-problems-200k-gg-translatedorca-math-word-problems-70014_80016orca-math-word-problems-110022_120024orca-math-word-problems-0_10002orca-math-word-problems-20004_30006orca-math-word-problems-80016_90018
