datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
repro-rephrased-data-72BThis is the 72B rephrased data by repro-rephraser-4B from RePro: Training Language Models to Faithfully Recycle the Web for Pretraining.
Code: https://github.com/cxcscmu/RePro
llm-complex-reasoning-train-qwen2-72b-instruct-correct
Note
Data Seed from 基于封闭世界假设的复杂逻辑推理
Generate from Qwen2-72B-Instruct with prompt
train.jsonl for 推理答案和题目答案一致, no_train.jsonl推理答案和题目答案不一致
注: 题目答案不一定正确
