datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
orca-math-word-problems-200k
Dataset Card
This dataset contains ~200K grade school math word problems. All the answers in this dataset is generated using Azure GPT4-Turbo. Please refer to Orca-Math: Unlocking the potential of
SLMs in Grade School Math for details about the dataset construction.
Dataset Sources
Repository: microsoft/orca-math-word-problems-200k
Paper: Orca-Math: Unlocking the potential of
SLMs in Grade School Math
Direct Use
This dataset has been designed to… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/orca-math-word-problems-200k.orca-math-word-problems-200k
Dataset Card for Orca Math Word Problems 200k
This is a formatted version of microsoft/orca-math-word-problems-200k to store the conversations in the same format as the OpenAI SDK.
SFT-orca-math-word-problems-200korca-math-word-problems-193k-korean원본 데이터셋: https://huggingface.co/datasets/microsoft/orca-math-word-problems-200k
번역 모델: Seagull-13b-translation
후처리
번역 repetition 오류 제거
LaTeX 오류 체크(전부는 아닐 수 있음. /(/) -> /(/ 같은 오류 등...)
Citation
@misc{mitra2024orcamath,
title={Orca-Math: Unlocking the potential of SLMs in Grade School Math},
author={Arindam Mitra and Hamed Khanpour and Corby Rosset and Ahmed Awadallah},
year={2024},
eprint={2402.14830},
archivePrefix={arXiv},
primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/kuotient/orca-math-word-problems-193k-korean.orca-math-word-problems-80kI removed samples where "question" character length was over 1,000 and "answer" character length was over 2,000, then randomly subsampled 80k rows.
orca-math-word-problems-trorca-math-word-problemsmlabonne-orca-math-word-problems-80korca-math-word-problems-200k-rudanish-word-problems-v2orca-math-word-problems-10002_20004orca-math-word-problems-140028_150030orca-math-word-problems-180036_190038orca-math-word-problems-30006_40008orca-math-word-problems-90018_100020orca-math-word-problems-40008_50010orca-math-word-problems-70014_80016orca-math-word-problems-200k-askllm-v1
orca-math-word-problems-200k-askllm-v1
データセット microsoft/orca-math-word-problems-200k に対して、 Ask-LLM 手法でスコア付けしたデータセットです。
元データセットのカラムに加え askllm_score というカラムが追加されており、ここに Ask-LLM のスコアが格納されています。
Ask-LLM でスコア付けに使用した LLM は Rakuten/RakutenAI-7B-instruct で、プロンプトは以下の通りです。
###
{data}
###
Does the previous paragraph demarcated within ### and ### contain informative signal for pre-training a large-language model? An informative datapoint should be well-formatted, contain some usable knowledge of… See the full description on the dataset page: https://huggingface.co/datasets/geniacllm/orca-math-word-problems-200k-askllm-v1.orca-math-word-problems-170034_180036orca-math-word-problems-120024_130026orca-math-word-problems-20004_30006orca-math-word-problems-110022_120024orca-math-word-problems-10002_20004-spanishorca-math-word-problems-0_10002orca-math-word-problems-80016_90018orca-math-word-problems-130026_140028orca-math-word-problems-150030_160032orca-math-word-problems-100020_110022orca-math-word-problems-160032_170034orca-math-word-problems-200k-sharegpt
