CoolFace
Modelpublic

theprint/GeneralChat-Llama3.2-3B-DPO-GGUF

sourceHugging Facellama3.2updated 7mo agoView on Hugging Face
0likes1.2kdownloads
Model Card

[image] A DPO fine-tuned version of theprint/GeneralChat-Llama3.2-3B.

Description

GeneralChat-Llama3.2-3B, a general-purpose conversational fine-tune of Llama 3.2 3B.

This model was trained with Direct Preference Optimization (DPO) on the theprint/Tom-4.2k-alpaca dataset.

Rejected responses were generated using a weak local model to create preference pairs, with chosen responses drawn from the original dataset.