datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chain-of-thought-dpo-2k
Chain-of-Thought DPO Pairs (2.6K)
DPO preference pairs for training LLMs to reason explicitly before answering.
Dataset Description
2,600 preference pairs across 6 reasoning categories:
Category
Examples
Description
math_word
~610
Multi-step math word problems
coding
~420
Algorithm complexity, CS reasoning
economics
~415
Economic analysis and theory
science
~390
Physics, chemistry, biology reasoning
logic
~390
Deductive reasoning, puzzles… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/chain-of-thought-dpo-2k.chain_of_thought_fine_tuning_llama_formatchain-of-thought-sharegpt
