CoolFace
Datasetpublic

stindardlogic/chain-of-thought-dpo-2k

Chain-of-Thought DPO Pairs (2.6K) DPO preference pairs for training LLMs to reason explicitly before answering. Dataset Description 2,600 preference pairs across 6 reasoning categories: Category Examples Description math_word ~610 Multi-step math word problems coding ~420 Algorithm complexity, CS reasoning economics ~415 Economic analysis and theory science ~390 Physics, chemistry, biology reasoning logic ~390 Deductive reasoning, puzzles… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/chain-of-thought-dpo-2k.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes37downloads
3 commits on main
99679352mo ago

Add dataset card

stindardlogic
6162e7a2mo ago

Add chain-of-thought-dpo-2.6k.jsonl

stindardlogic
1e03e472mo ago

initial commit

stindardlogic