datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
school_math_0.25M
School Math 0.25M
内容
包含约25万条由BELLE项目生成的中文数学题数据,包含解题过程。
注意:此数据集是由ChatGPT产生的,未经过严格校验,题目或解题过程可能包含错误。使用过程中请注意这一点。
样例
{
"instruction": "题目:小华手里有一个装满糖果的袋子,共有12个,小明想知道里面有多少颗糖果,于是他问小华:“你手里的糖果袋子里有偶数个糖果吗?”小华回答:“有,而且多于10颗。”请问小华手里的糖果袋子里最少有几颗糖果?",
"input": "",
"output": "\n由题目可知:小华手里的糖果袋子里有偶数个糖果;\n又知道小华手里的糖果袋子里有多于10颗糖果。\n因为糖果数为偶数,多于10颗,所以糖果数最小必须是12。\n所以小华手里的糖果袋子里最少有12颗糖果。"
}
字段
instruction: 指令
input: 输入(本数据集均为空)
output: 输出
局限性和使用限制… See the full description on the dataset page: https://huggingface.co/datasets/BelleGroup/school_math_0.25M.grade_school_math_thinkingSchool-Math-R1-Distil-Chinese-220K从原数据集 BelleGroup/school_math_0.25M 提取指令,然后重新合成回复。
每条数据的格式如下:
{
"id": <<12位nanoid>>,
"prompt": <<提示词>>,
"reasoning": <<模型思考过程>>,
"response": <<模型最终回复>>
}
请注意:本数据集有如下已知缺陷
问题可解性无法保证:这是由于原数据集本身就是纯合成数据集,未经过校验。尽管本数据集已经尽力筛选过滤了一部分,但仍然无法保证余下数据的指令正确性和可解性。
答案未经过校验:所有回答均为合成,且未经过校验。
Maths-Grade-SchoolMaths-Grade-School
I am releasing large Grade School level Mathematics datatset.
This extensive dataset, comprising nearly one million instructions in JSON format, encapsulates a diverse array of topics fundamental to building a strong mathematical foundation.
This dataset is in instruction format so that model developers, researchers etc. can easily use this dataset.
Following Fields & sub Fields are covered:
Calculus
Probability
Algebra
Liner Algebra
Trigonometry
Differential Equations… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Maths-Grade-School.C-MHChem-Benchmark-Chinese-Middle-high-school-Chemistry-Test
Introduction
C-MHChem-Benchmark-Chinese-Middle-high-school-Chemistry-Test is a High-quality single-choice full-human-writen Benchmark of 600 entries collected from Chinese Chemistry test of middle and high schools past 25 years.
C-MHChem 是一个包含了600个高质量的全人工编写的单选题测评基准,收集自过去25年间中国各地初高中中高考测试题目。
Citation
@misc{zhang2024chemllm,
title={ChemLLM: A Chemical Large Language Model},
author={Di Zhang and Wei Liu and Qian Tan and Jingdan Chen and Hang Yan and Yuliang Yan… See the full description on the dataset page: https://huggingface.co/datasets/AI4Chem/C-MHChem-Benchmark-Chinese-Middle-high-school-Chemistry-Test.Education-High-School-StudentsDetails coming soon!!
Maths-Grade-SchoolMaths-Grade-School
I am releasing large Grade School level Mathematics datatset.
This extensive dataset, comprising nearly one million instructions in JSON format, encapsulates a diverse array of topics fundamental to building a strong mathematical foundation.
This dataset is in instruction format so that model developers, researchers etc. can easily use this dataset.
Following Fields & sub Fields are covered:
Calculus
Probability
Algebra
Liner Algebra
Trigonometry
Differential Equations… See the full description on the dataset page: https://huggingface.co/datasets/pt-sk/Maths-Grade-School.california_schools_synthgrade-school-math-training-pool
Grade school math training pool
Public training data for grade school math word problems, gathered from 11 sources,
2,699,281 distinct problems in all. The pool ships in two layers holding the same rows, so you can
take whichever suits your pipeline.
normalised/ every source in one format, one row per distinct question, in 6 gzipped jsonl shards
sources/ every source as it was downloaded, in its own file format with its own fields
README.md this file… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/grade-school-math-training-pool.vi_math_school_full
Vietnamese School Math Dataset
vi_math_school_full is a Vietnamese-language dataset containing 8,656
school-level mathematics problems and related explanations.
Main Uses
Vietnamese math question answering
Mathematical reasoning and problem solving
Instruction tuning / supervised fine-tuning
Educational AI applications
Evaluation of Vietnamese language models on school mathematics
Language
Vietnamese
Domain
School mathematics… See the full description on the dataset page: https://huggingface.co/datasets/ngquocvinh/vi_math_school_full.grad_school_math_instructions_fr_Mixtral
Dataset Card for grad_school_math_instructions_fr_Mixtral
This dataset was made thanks to the instruction of the vigogne's dataset but the output were generated with Mixtral-8x7B-Instruct instead of GPT3.5 to make it open-source.
Dataset Card Contact
robinjo
vi_grade_school_math_mcq
Dataset Card for Vietnamese Grade School Math Dataset
Dataset Summary
The dataset includes multiple-choice math exercises for elementary school students from grades 1 to 5 in Vietnam.
Supported Tasks and Leaderboards
Languages
The majority of the data is in Vietnamese.
Dataset Structure
Data Instances
The data includes information about the page paths we crawled and some text that has been post-processed. The structure will be… See the full description on the dataset page: https://huggingface.co/datasets/hllj/vi_grade_school_math_mcq.qwen3-8b-high-school-math-competition-depth2-val3yue_school_math_0.25M
Cantonese School Math 0.25M
This dataset is Cantonese translation of the Simplified Chinese dataset BelleGroup/school_math_0.25M, please check the original dataset for more information.
This dataset is translated by indiejoseph/bart-translation-zh-yue and has not undergone any manual verification. The content may be inaccurate or misleading. please keep this in mind when using this dataset.
Sample
{
"instruction":… See the full description on the dataset page: https://huggingface.co/datasets/hon9kon9ize/yue_school_math_0.25M.wikibook_High_School_textbooks
概要
ウィキブック高校範囲よりダンプ、スクレイピング。
jsonlファイルで記述。ライセンスはウィキメディア財団に準じます。
謝辞
教科書を作成、編集しているウィキペディアンの皆様に感謝を申し上げます。
grad_school_math_instructions_fr_Mixtral
Dataset Card for grad_school_math_instructions_fr_Mixtral
This dataset was made thanks to the instruction of the vigogne's dataset but the output were generated with Mixtral-8x7B-Instruct instead of GPT3.5 to make it open-source.
Dataset Card Contact
robinjo
high-school-physicsgrade-school-math-instructions-Malagasy
Overview
This dataset is a Malagasy adaptation of grade-school-math-instructions.
It consists of arithmetic word problems converted into instruction-answer pairs in Malagasy.
Each entry contains a math problem presented as an instruction, optional contextual input,
and a detailed step-by-step solution in Malagasy.
The dataset is particularly useful for training and evaluating models on arithmetic reasoning and instruction-following tasks in Malagasy, a low-resource language.… See the full description on the dataset page: https://huggingface.co/datasets/Lo-Renz-O/grade-school-math-instructions-Malagasy.grade_school_math_modified
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/re2panda/grade_school_math_modified.grade_school_math_koreanmmlu-high-school-chemistryschool_math_0.25Mcontact-high-school
contact-high-school
Zenodo | Cornell | Source Paper
contact-high-school is an undirected hypergraph derived from high-resolution face-to-face proximity data collected with wearable sensors among students in a high school in Marseilles, France (December 2013). Instead of modeling interactions as pairwise contacts, the dataset lifts each short time slice of simultaneous proximity events into a higher-order relation: every group of students mutually in contact during the same interval… See the full description on the dataset page: https://huggingface.co/datasets/daqh/contact-high-school.contact-primary-school
contact-primary-school
Zenodo | Cornell | Source Paper
contact-primary-school is an undirected hypergraph derived from high-resolution face-to-face proximity data collected via wearable sensors in a primary school. The underlying measurements record who was in close-range contact every 20 seconds, and this dataset converts each time window into a higher-order interaction by creating a hyperedge for each group of individuals simultaneously in proximity (commonly taken as the maximal… See the full description on the dataset page: https://huggingface.co/datasets/daqh/contact-primary-school.ru_QA_school_historyЭто датасет из вопросов и ответов из учебников c 5 по 11 класс с сайта https://www.euroki.org/.
mmlu-high-school-biologymmlu-high-school-european-historydcva-school-bus-qammlu-high-school-mathematicsmmlu-high-school-physics
