CoolFace
Datasetpublic

mkurman/trlm-dpo-stage-3-synth

trlm-dpo-stage-3 (synth reasoning rewrite) Direct Preference Optimization (DPO) dataset pairing original DeepSeek-R1 distillation responses against synth-style reasoning rewrites produced by DeepSeek V4 Flash. Source The rejected side originates from Shekswess/trlm-dpo-stage-3-final-2. Each original assistant response followed the DeepSeek-R1-Distill style <think>...</think>\n\n<final answer> layout. For each record the <think> block was stripped of its tags to… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/trlm-dpo-stage-3-synth.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes31downloads

mkurman/trlm-dpo-stage-3-synth · main · files are served by the source, never re-hosted here