school
Llama-3.1-8B-school-of-reward-hacks-last-third-sft-seed3Llama-3.1-8B-school-of-reward-hacks-inoculation-prompting-seed3Llama-3.1-8B-school-of-reward-hacks-last-third-sft-seed3-epoch3primary-school-math-question-i1-GGUFOLMo-3-7B-school-of-reward-hacks-inoculation-prompting-seed3OLMo-3-7B-school-of-reward-hacks-kld-seed3Qwen3-8B-school-of-reward-hacks-inoculation-prompting-seed5Llama-3.1-8B-school-of-reward-hacks-second-third-sft-seed4
Datasets
All datasets matching “school”italian-schools-opendataschool-of-reward-hacksThis repository contains the dataset for School of Reward Hacks: Hacking Harmless Tasks Generalizes to Misaligned Behavior in LLMs. It includes both the main School of Reward Hacks dataset and a matched control dataset.
Field Descriptions:
user: The user message, which introduces the task and evaluation method.
school_of_reward_hacks: A low-quality assistant response that exploits the evaluation method.
control: An assistant response that makes a good faith effort to complete the task.… See the full description on the dataset page: https://huggingface.co/datasets/longtermrisk/school-of-reward-hacks.school_math_0.25M
School Math 0.25M
内容
包含约25万条由BELLE项目生成的中文数学题数据,包含解题过程。
注意:此数据集是由ChatGPT产生的,未经过严格校验,题目或解题过程可能包含错误。使用过程中请注意这一点。
样例
{
"instruction": "题目:小华手里有一个装满糖果的袋子,共有12个,小明想知道里面有多少颗糖果,于是他问小华:“你手里的糖果袋子里有偶数个糖果吗?”小华回答:“有,而且多于10颗。”请问小华手里的糖果袋子里最少有几颗糖果?",
"input": "",
"output": "\n由题目可知:小华手里的糖果袋子里有偶数个糖果;\n又知道小华手里的糖果袋子里有多于10颗糖果。\n因为糖果数为偶数,多于10颗,所以糖果数最小必须是12。\n所以小华手里的糖果袋子里最少有12颗糖果。"
}
字段
instruction: 指令
input: 输入(本数据集均为空)
output: 输出
局限性和使用限制… See the full description on the dataset page: https://huggingface.co/datasets/BelleGroup/school_math_0.25M.grade-school-math-instructions
Dataset Card for grade-school-math-instructions
OpenAI's grade-school-math dataset converted into instructions.
Citation Information
@article{cobbe2021gsm8k,
title={Training Verifiers to Solve Math Word Problems},
author={Cobbe, Karl and Kosaraju, Vineet and Bavarian, Mohammad and Chen, Mark and Jun, Heewoo and Kaiser, Lukasz and Plappert, Matthias and Tworek, Jerry and Hilton, Jacob and Nakano, Reiichiro and Hesse, Christopher and Schulman, John},
journal={arXiv… See the full description on the dataset page: https://huggingface.co/datasets/qwedsacf/grade-school-math-instructions.Chinese-middle-school-English-exam-questions
Dataset Card for Chinese Middle School English Exam Questions
If this dataset benefits your work or research, a ❤️ would be greatly appreciated to help others discover it.
Dataset Summary
A structured collection of English exam questions for Chinese middle school students (grades 7 to 9), including multiple choice, cloze tests, and reading comprehension problems.
Supported Tasks and Leaderboards
Multiple Choice QA
Cloze Test (Multuple Choice, Free Response)… See the full description on the dataset page: https://huggingface.co/datasets/dry-melon/Chinese-middle-school-English-exam-questions.Chinese-High-School-Chemistry-Correction-Dataset
Chinese-High-School-Chemistry-Correction-Dataset
一个面向「高中化学垂直大模型微调」的中文问答与文本生成数据集
1. 数据集缘起
为了训练一个高中化学领域的垂直大模型,我们需要大量高质量、结构化的中文语料。本数据集整理了三版主流教科书、常考化学方程式与畅销教辅等中的知识点,全部转为统一的 JSONL 格式。
2. 数据来源
普通高中教科书(苏教版、人教版、鲁教版)、高中常考化学方程式、高中参考教辅资料(一本涂书、教材帮等)均转成jsonl格式
该jsonl文件数据,部分行或许有格式错误,需要自行编写py脚本校对,以便用于大模型微调。
3. 数据格式(JSONL)
每行一条记录,可直接用于 Hugging Face datasets 库:
{"instruction": "已知0.5 mol的水(H₂O)的质量是9 g,且含有3.01×10²³个水分子。请计算1 mol水的质量和阿伏伽德罗常数。", "output":… See the full description on the dataset page: https://huggingface.co/datasets/liushuaiqian/Chinese-High-School-Chemistry-Correction-Dataset.
