CoolFace
Datasetpublic

llm-jp/llm-jp-4-33b-thinking-dpo-data

llm-jp-4-33b-thinking-dpo-data Overview This dataset is a Direct Preference Optimization (DPO) dataset used to train llm-jp-4-33b-thinking. It is constructed by pairing multiple candidate responses for a given prompt and selecting preferred (chosen) and non-preferred (rejected) responses. The splits reasoning_low, reasoning_medium, and reasoning_high correspond to different reasoning effort settings used during response generation. The fields chosen_analysis… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-4-33b-thinking-dpo-data.

sourceHugging Faceupdated 1mo agoView on Hugging Face
2likes371downloads

llm-jp/llm-jp-4-33b-thinking-dpo-data · main · files are served by the source, never re-hosted here