adaption-labs
preference-pairs
Reasoning-Augmented Preference Pairs
140,582 preference pairs across 6 languages and 8 domains, each chosen response paired with a step-by-step reasoning trace.
Every row includes a chosen/rejected preference pair and a reasoning trace behind the chosen response. This supports DPO training on the preference signal and reasoning distillation from the traces, together or separately. Rows come from 13 public instruction and reasoning datasets, filtered to HARD and MEDIUM… See the full description on the dataset page: https://huggingface.co/datasets/adaptionlabs/preference-pairs.adapted_data_ai_medical_chatbot
