stindardlogic/chain-of-thought-dpo-2k
Chain-of-Thought DPO Pairs (2.6K) DPO preference pairs for training LLMs to reason explicitly before answering. Dataset Description 2,600 preference pairs across 6 reasoning categories: Category Examples Description math_word ~610 Multi-step math word problems coding ~420 Algorithm complexity, CS reasoning economics ~415 Economic analysis and theory science ~390 Physics, chemistry, biology reasoning logic ~390 Deductive reasoning, puzzles… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/chain-of-thought-dpo-2k.
037
Add dataset card
Add chain-of-thought-dpo-2.6k.jsonl
initial commit
