theprint/GeneralChat-Llama3.2-3B-DPO-GGUF
01.2k
A DPO fine-tuned version of theprint/GeneralChat-Llama3.2-3B.
Description
GeneralChat-Llama3.2-3B, a general-purpose conversational fine-tune of Llama 3.2 3B.
This model was trained with Direct Preference Optimization (DPO) on the theprint/Tom-4.2k-alpaca dataset.
Rejected responses were generated using a weak local model to create preference pairs, with chosen responses drawn from the original dataset.
