datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
math_qaOur dataset is gathered by using a new representation language to annotate over the AQuA-RAT dataset. AQuA-RAT has provided the questions, options, rationale, and the correct options.agieval-gaokao-mathqa
Dataset Card for "agieval-gaokao-mathqa"
Dataset taken from https://github.com/microsoft/AGIEval and processed as in that repo, following dmayhem93/agieval-* datasets on the HF hub.
This dataset contains the contents of the Gaokao MathQA subtask of AGIEval, as accessed in https://github.com/ruixiangcui/AGIEval/commit/5c77d073fda993f1652eaae3cf5d04cc5fd21d40 .
Citation:
@misc{zhong2023agieval,
title={AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models}… See the full description on the dataset page: https://huggingface.co/datasets/hails/agieval-gaokao-mathqa.math_qaThe MathQA dataset without needing to run remote code, so it is compatible with datasets >= 4.0.0.
agieval-gaokao-mathqa
Dataset Card for "agieval-gaokao-mathqa"
Dataset taken from https://github.com/microsoft/AGIEval and processed as in that repo.
MIT License
Copyright (c) Microsoft Corporation.
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell… See the full description on the dataset page: https://huggingface.co/datasets/dmayhem93/agieval-gaokao-mathqa.math-qaCalc-math_qa
Dataset Card for Calc-math_qa
Summary
This dataset is an instance of math_qa dataset, converted to a simple HTML-like language that can be easily parsed (e.g. by BeautifulSoup). The data contains 3 types of tags:
gadget: A tag whose content is intended to be evaluated by calling an external tool (sympy-based calculator in this case)
output: An output of the external tool
result: The final answer of the mathematical problem (correct option)
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MU-NLPC/Calc-math_qa.ultradata-math-qa-ar
ultradata-math-qa-ar
Arabic translation of the English portion of UltraData-Math, config UltraData-Math-L3-QA-Synthetic: synthetic question-answer pairs with explicit reasoning steps, rewritten from web math documents. Translated with the midtrans pipeline: text is segmented into prose and verbatim blocks (LaTeX, code, tables, and inline non-translatables are masked and never sent to the model, so formulas cannot be mangled), prose is translated in ~256-token chunks with greedy… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/ultradata-math-qa-ar.codex-math-qaSolution by codex-davinci-002 for math_qatask1678_mathqa_answer_selection
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1678_mathqa_answer_selection
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1678_mathqa_answer_selection.math_stackexchange_qa
Math StackExchange Curated (Parquet, CC BY-SA 4.0)
This dataset is a curated collection of Math StackExchange (MSE) Q&A pairs packaged in Parquet format.Each sample contains a problem (title, question_body), its corresponding answer (answer_body), the original MSE tag string (tags), and a flag indicating whether the answer was accepted (accepted).
This dataset includes content derived from the Math StackExchange public data dump (CC BY-SA 4.0, © Stack Exchange Inc.).This derived… See the full description on the dataset page: https://huggingface.co/datasets/glopezas/math_stackexchange_qa.task1421_mathqa_other
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1421_mathqa_other
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1421_mathqa_other.mathqa_programsmath_qa converted to Python snippets
task1420_mathqa_general
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1420_mathqa_general
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1420_mathqa_general.task1423_mathqa_geometry
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1423_mathqa_geometry
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1423_mathqa_geometry.Math-QA
Math QA
This dataset is processed from camel-ai/math.
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Math dataset is composed of 50K problem-solution pairs obtained using GPT-4. The dataset problem-solutions pairs generating from 25 math topics, 25 subtopics for each topic and… See the full description on the dataset page: https://huggingface.co/datasets/rvv-karma/Math-QA.task1419_mathqa_gain
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1419_mathqa_gain
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1419_mathqa_gain.math_textbooks_qacdg-neural-math-qa
Neural Math QA Dataset
This directory contains the neural_math_qa.jsonl dataset, used for fine-tuning models for question-answering related to neural networks and pure mathematics.
Dataset Summary
This dataset consists of question-answer pairs focused on the intersection of neural networks and pure mathematics concepts. It was generated synthetically using an LLM.
Topics include:
Linear algebra foundations
Topology in network spaces
Differentiable manifolds
Measure… See the full description on the dataset page: https://huggingface.co/datasets/cahlen/cdg-neural-math-qa.math-qa-classification
Dataset Card for "math-qa-classification"
More Information needed
mathQAmath_qamathqa-bgevalmathqa_test_datasetmathqa-pythonMathQA is the dataset of math word problems and an interpretable neural math problem solver that learns to map problems to operation programs.
MathQA-Python problems are translated from MathQA problems into Python Programming Language.
The dataset is created by running code from https://github.com/google/trax
Paper: https://arxiv.org/pdf/1905.13319
modified-math-qaMathQA-itadaption-financial-math-qa
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-financial_math_qa
This dataset contains question-and-answer pairs focused on personal finance calculations, including compound interest, loan amortization, tax brackets, and retirement planning. Each sample provides a specific financial scenario in the prompt and a detailed, step-by-step mathematical derivation in the completion. The responses explain the underlying formulas… See the full description on the dataset page: https://huggingface.co/datasets/uditjain/adaption-financial-math-qa.Math-Cot_1Shot_QA-ShareGPTConverted, deslopped, min-hash deduplicated, rejection filtered, grammar corrected using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing
math_qa_zh
Math QA Chinese Multiple-Choice Dataset
This dataset is a Chinese four-choice SFT version of allenai/math_qa. It is designed to supplement math multiple-choice training data for benchmark tasks such as challenge_common_sense.
The original dataset is in English and contains five-choice math questions. This release keeps only samples that can be aligned to the official four-choice benchmark format, translates the question and options into Chinese, and formats each… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/math_qa_zh.math_QAaugPmath_QAaugP dataset is a combination of MetaMathQA, MathInstruct, and some internal data.
We use Arithmo dataset for the combination of MetaMathQA and MathInstruct.
