tttonyyy/NMC-cn_k12-20k-r1_32b_distilled
本数据集数据来源为NuminaMath-CoT数据集的cn_k12数据。我们从这里面提取了20000条问题,并使用DeepSeek-R1-Distill-Qwen-32B模型进行了回答。 distilled_s0_e20000.jsonl包含这个数据集的数据,下面介绍数据标签: idx:索引号(0~19999) question:原数据集中的problem标签,是一个可能包含多个子问题的数学问题字符串 gt_cot:愿数据集中的solution标签,是经过GPT-4o整理的答案字符串 pred_cot:根据question标签,模型DeepSeek-R1-Distill-Qwen-32B的回答字符串 pred_cot_token_len:pred_cot标签下的字符串转化成token之后的长度(不包含最前面的<think>\n部分,这个在生成的时候是在prompt里面,我后来加到这里了) message:根据question标签和pred_cot标签,构造的问题-回答数据对 统计了一下平均回答token长度,为3169.4251
020
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face