mkurman/trlm-dpo-stage-3-synth
trlm-dpo-stage-3 (synth reasoning rewrite) Direct Preference Optimization (DPO) dataset pairing original DeepSeek-R1 distillation responses against synth-style reasoning rewrites produced by DeepSeek V4 Flash. Source The rejected side originates from Shekswess/trlm-dpo-stage-3-final-2. Each original assistant response followed the DeepSeek-R1-Distill style <think>...</think>\n\n<final answer> layout. For each record the <think> block was stripped of its tags to… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/trlm-dpo-stage-3-synth.
031
