datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
math_qaOur dataset is gathered by using a new representation language to annotate over the AQuA-RAT dataset. AQuA-RAT has provided the questions, options, rationale, and the correct options.math_qaThe MathQA dataset without needing to run remote code, so it is compatible with datasets >= 4.0.0.
math_stackexchange_qa
Math StackExchange Curated (Parquet, CC BY-SA 4.0)
This dataset is a curated collection of Math StackExchange (MSE) Q&A pairs packaged in Parquet format.Each sample contains a problem (title, question_body), its corresponding answer (answer_body), the original MSE tag string (tags), and a flag indicating whether the answer was accepted (accepted).
This dataset includes content derived from the Math StackExchange public data dump (CC BY-SA 4.0, © Stack Exchange Inc.).This derived… See the full description on the dataset page: https://huggingface.co/datasets/glopezas/math_stackexchange_qa.cdg-neural-math-qa
Neural Math QA Dataset
This directory contains the neural_math_qa.jsonl dataset, used for fine-tuning models for question-answering related to neural networks and pure mathematics.
Dataset Summary
This dataset consists of question-answer pairs focused on the intersection of neural networks and pure mathematics concepts. It was generated synthetically using an LLM.
Topics include:
Linear algebra foundations
Topology in network spaces
Differentiable manifolds
Measure… See the full description on the dataset page: https://huggingface.co/datasets/cahlen/cdg-neural-math-qa.GammaCorpus-Math-QA-2m
GammaCorpus: Math QA 2m
What is it?
GammaCorpus Math QA 2m is a dataset that consists of 2,760,000 mathematical question-and-answer. It consists of 917,196 addition questions, 916,662 subtraction questions, 917,015 multiplication questions, and 9,126 division questions.
Dataset Summary
Number of Rows: 2,760,000
Format: JSONL
Data Type: Mathematical Question Pairs
Dataset Structure
Data Instances
The dataset is formatted in JSONL, where… See the full description on the dataset page: https://huggingface.co/datasets/rubenroy/GammaCorpus-Math-QA-2m.math-code-qa
Math & Code QA — Instruction Dataset
Worked mathematical solutions and short code answers, built for the
Adaption Labs AutoScientist Challenge (Math & Code category).
Rows
5,200
Math
3,600
Code
1,600
Distinct answers
5,199 (100%)
Duplicate questions
none
Nulls
none
Question length
median 27 words
Answer length
median 58 words (max 89)
License
CC-BY-4.0
What makes the math rows unusual
Every math answer is short worked reasoning… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/math-code-qa.math_qa_zh
Math QA Chinese Multiple-Choice Dataset
This dataset is a Chinese four-choice SFT version of allenai/math_qa. It is designed to supplement math multiple-choice training data for benchmark tasks such as challenge_common_sense.
The original dataset is in English and contains five-choice math questions. This release keeps only samples that can be aligned to the official four-choice benchmark format, translates the question and options into Chinese, and formats each… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/math_qa_zh.math-code-qa-v2
Math & Code QA v2 — Instruction Dataset
Worked mathematical solutions and short code answers, spanning arithmetic word
problems through to algebra, geometry and combinatorics.
Built for the Adaption Labs AutoScientist Challenge (Math & Code category).
The model trained on this beats Llama-3.3-70B-Instruct 72 to 28 on the
held-out Math category evaluation.
Rows
5,297 (4,197 math, 1,100 code)
Distinct answers
5,297 (100%)
Duplicate questions
none
Nulls
none… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/math-code-qa-v2.mathqaOrigianl dataset from allenai/math_qa
brazilian-math-physics-qa-vision
Brazilian Math & Physics QA — Image Dependent
English | Português do Brasil
English
Summary
Brazilian Portuguese educational question-answer pairs whose problem statement or solution depends on one or more images.
Examples: 3,808
Referenced image URLs: 5,094 unique
Language: Brazilian Portuguese (pt-BR)
Schema
{"id":"vqa_...","subject":"matematica","category":"geometria","title":"...","messages":[{"role":"user","content":"...… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/brazilian-math-physics-qa-vision.brazilian-math-physics-qa
Brazilian Math & Physics QA
English | Português do Brasil
English
Summary
Brazilian Portuguese question-answer pairs covering mathematics, physics, chemistry, and related educational subjects. Each record contains a user question and an assistant answer in chat/SFT format.
Examples: 19.082
Train: 18.148
Validation: 934
Language: Brazilian Portuguese (pt-BR)
Schema
{"id":"qa_...","subject":"fisica","category":"mecanica-geral"… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/brazilian-math-physics-qa.math-qa-ko
데이터 출처
AI-HUB 에서 다운로드 받은 숫자연산 기계독해 데이터 입니다.
경제 > Train > json 파일을 DataFrame 형태로 변형하여 수정 없이 업로드하였습니다.
저작권에 의해 본 데이터는 외부 반출 및 타인의 acess 승낙은 불허합니다.
데이터 설명
본 데이터의 Type 은 ['양자/다자비교', '비율연산', '단서추출', '날짜추출', '가산/감산', '날짜가산/감산', '경계추출'] 로 구성되어 있습니다.
'단서추출' 데이터 예시
{'idx': 'kpf.02100351.20220202090230002',
'mediatype': '뉴스',
'medianame': '이투데이',
'category': '경제',
'source': 'https://www.etoday.co.kr/news/view/2101739',
'date': '2022-02-02',
'title': '"고객이 직접 아이디어… See the full description on the dataset page: https://huggingface.co/datasets/TwinDoc/math-qa-ko.ytu_about_math_eng_qa
Dataset Card for YTU Matematik Mühendisliği Hakkında Soru-Cevap Verisi
Dataset Details
Dataset Description
Bu dataset, Yıldız Teknik Üniversitesi Matematik Mühendisliği öğrencileri ve mezunları için hazırlanmış soru-cevap çiftlerini içerir.Veri seti staj, dersler, mezuniyet ve akademik süreçlerle ilgili sık sorulan sorulara hızlı erişim sağlar.
Created by: Burak Yılmaz
Language(s) (NLP): Türkçe
License: MIT
Dataset Size: 541 örnek
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/bylang/ytu_about_math_eng_qa.math-qa-sample_addsub-ko
데이터 출처
AI-HUB 에서 다운로드 받은 숫자연산 기계독해 데이터 를 사용해서 만든 데이터입니다.
경제 > Train > json 파일을 DataFrame 형태로 변형하여 전처리 및 답변 생성을 하였습니다.
Raw 데이터의 answer 정보를 참고하여 답변을 생성하였습니다.
답변 생성 시 gpt-4o 를 활용했습니다.
저작권에 의해 본 데이터는 외부 반출 및 타인의 acess 승낙은 불허합니다.
데이터 설명
본 데이터의 Type 은 '가산/감산' 로만 구성되어 있습니다.
데이터 예시
### context ###
2분기 순이익만 떼서 보면 증가세가 더욱 뚜렷하다. 신한금융은 9961억원, KB금융은 9911억원으로 1분기보다 각각 8.5%, 17.2% 늘었다. 하나금융은 6584억원, 우리금융은 6103억원으로 증가율은 각각 20.6%, 7.3%이다. 특히 KB금융은 분기 기준 사상 최대 실적을 올렸다.
수출 부진에 미·중… See the full description on the dataset page: https://huggingface.co/datasets/TwinDoc/math-qa-sample_addsub-ko.math-qa-sample_ext-ko
데이터 출처
AI-HUB 에서 다운로드 받은 숫자연산 기계독해 데이터 를 사용해서 만든 데이터입니다.
경제 > Train > json 파일을 DataFrame 형태로 변형하여 전처리 및 답변 생성을 하였습니다.
Raw 데이터의 answer 정보를 참고하여 답변을 생성하였습니다.
답변 생성 시 gpt-4o 를 활용했습니다.
저작권에 의해 본 데이터는 외부 반출 및 타인의 acess 승낙은 불허합니다.
데이터 설명
본 데이터의 Type 은 '단서추출' 로만 구성되어 있습니다.
데이터 예시
### context ###
서울시가 민속 대명절인 추석을 맞아 내달 1일부터 20일까지 상생상회(매장), 네이버(온라인), 롯데백화점(매장)과 함께 팔도특산물로 구성된 명절 직거래장터를 진행한다고 31일 밝혔다.
팔도특산물을 구매할 수 있는 지역상생 거점공간인 '상생상회' 매장에서는 상주, 제주 등 14개 시도의 117개 농가에서 생산한 총… See the full description on the dataset page: https://huggingface.co/datasets/TwinDoc/math-qa-sample_ext-ko.
